Project Context: Sun Country Airlines, a Minneapolis-based low-cost leisure carrier operating out of MSP (Minneapolis-Saint Paul International Airport), had no dedicated data analyst on staff and relied entirely on anecdotal assumptions to drive marketing decisions. Acting as an analytics consultant for a BANA 200 case study, I applied K-Means clustering, a machine learning technique that groups similar customers together, to 15,144 customer-trip records spanning 2013 to 2014, identifying five actionable segments and translating the findings into three targeted campaigns tied to leadership's stated goals around Ufly loyalty enrollment, direct channel growth, and vacation package differentiation.
Business Question: Which distinct customer segments exist within Sun Country's booking data, and what targeted strategies should the airline pursue to grow loyalty enrollment, increase direct bookings, and differentiate its vacation packages?
K-Means with k=5 revealed five distinct, business-interpretable groups. The clearest pattern: Clusters 0 and 1 share the same Minneapolis leisure traveler profile but split entirely on where they book, direct versus third-party platforms like Expedia. Loyalty enrollment was almost entirely absent outside Cluster 3.
Two numbers from the initial exploratory analysis (EDA) foreshadowed everything the clustering later confirmed: 78.5% of customers weren't enrolled in Ufly Rewards, Sun Country's loyalty program, and bookings split almost evenly between direct (45.9%, via SCA.com) and third-party channels (44.0%, Expedia and similar).
- Data prep: started with 15,144 records across 90 features, removed two identifier columns with no analytical signal, worked with pre-normalized data and no missing values.
- EDA: surfaced the near-even booking channel split and the 78.5% non-Ufly rate before any model ran, both of which shaped how every cluster was later interpreted.
- Choosing k=5: tested k=2 through 10 using the Elbow Method (plotting error against cluster count and looking for where more clusters stop helping) and Silhouette Score (how cleanly each customer fits its cluster), both pointed to k=5 as the best balance of statistical fit and business interpretability.
- The channel split between C0 and C1 was the most strategically interesting finding, two nearly identical traveler profiles that needed two completely different interventions.
- Naming segments mattered. Personas like "MSP Direct Bookers" made the results usable for stakeholders who'd never read a centroid table.
- Unsupervised learning runs on judgment, not accuracy scores. Deciding how many clusters to use and what a centroid means in business terms is where the real analytical work happens.