Predicting Employee Churn

UC Irvine · MSBA Coursework

09/2025-12/2025

Project Context: Employee turnover imposes substantial costs on organizations through lost expertise, productivity disruptions, and recruitment and training expenses. For this BANA 273 machine learning project, my team analyzed a publicly available Kaggle dataset of 4,653 employees across three major Indian cities to identify which employee characteristics most strongly predict voluntary churn, an employee leaving on their own, and to build a model capable of flagging at-risk employees early enough for targeted intervention.

Business Question: Which employee attributes and organizational factors most strongly influence churn, and how accurately can they be used to identify at-risk employees?

Who Is Most Likely to Leave?
  • Mid-pay tier employees churn the most, at 2.2x the odds of low-pay employees, sharpest when pay feels neither entry-level nor rewarding.
  • Master's degree holders churn 2.4x more than Bachelor's holders, likely reflecting greater external mobility and a mismatch with available advancement.
  • Benched employees, those without a current project assignment, churn 1.5x more than employees who have never been benched.
  • Pune employees churn 1.6x more than Bangalore employees, pointing to regional labor market or workplace differences.
  • Female employees show a higher churn rate than male employees, consistent across both descriptive analysis and the model.
Model Performance

Two models were compared on recall (the share of actual churners correctly caught), the priority metric since missing a churner costs far more than a false alarm.

Logistic Regression
63%
Recall (Churn Class)
Identified education, pay tier, city, benching, and gender as the strongest predictors, with strong interpretability through odds ratios.
Tuned Decision Tree
79%
Recall (Churn Class)
Joining year and pay tier were the dominant splits, capturing nonlinear interactions the logistic model can't.

The two models tell complementary stories: logistic regression surfaces linear trends and effect sizes, while the decision tree reveals how tenure, pay, and education combine into risk pathways neither variable alone would predict.

High Priority
Target Long-Tenured Mid-Pay Employees First
Primary risk group across both models
Joining year and pay tier were the two dominant predictors of churn. Concentrate outreach, career-development conversations, and pay adjustments on longer-tenured, mid-pay employees before turnover occurs.
High Priority
Review Mid-Tier Compensation Structure
Mid-pay tier · 19.7% of employees, 2.2x churn odds
Mid-tier dissatisfaction was the most consistent finding across both models. A pay equity analysis with clearer promotion pathways could close the gap between contribution and compensation driving this.
Medium Priority
Minimize Bench Time and Improve Communication During Gaps
Ever-benched employees · 10.3% of workforce, 1.5x churn odds
Reducing unassigned periods where possible, and proactively communicating during unavoidable bench time, can reduce the sense of exclusion it appears to create.
Medium Priority
Develop Location-Specific Retention Programs for Pune
Pune employees · 1.6x churn odds vs Bangalore
Retention strategy shouldn't be uniform across cities. Compensation benchmarks and engagement programs should be evaluated and tailored specifically for the Pune workforce.
Lower Priority
Create Advancement Pathways for Highly Educated Employees
Master's degree holders · 18.8% of workforce, 2.4x churn odds
Master's holders churn at more than double the rate of Bachelor's employees. Clearer recognition and advancement tracks can close the gap between expectation and reality driving this.
Methodology
  • Data prep: 4,653 records across 9 attributes from Kaggle, no missing values; categorical variables one-hot encoded for logistic regression, used directly by the decision tree; 70/30 train-test split, random state 42.
  • Class imbalance: 65.6% of employees stay, only 34.4% churn, so both models used balanced class weighting and recall became the primary evaluation metric instead of accuracy.
  • Modeling: logistic regression improved via GridSearchCV (L1 regularization) from 0.41 to 0.63 recall; the decision tree, pre-pruned and with its classification threshold (the probability cutoff for flagging a churner) lowered from 0.50 to 0.35, reached 0.79 recall. Both were confirmed stable via 5-fold cross-validation.
Takeaways
  • Choosing the right evaluation metric mattered more than choosing the right model. Once recall replaced accuracy, a model that looked decent was actually catching fewer than half the churners it needed to find.
  • Threshold tuning is an underused lever. Lowering the cutoff from 0.50 to 0.35 lifted the tree's recall from 0.64 to 0.79 with no additional feature engineering.
  • The clearest output of a churn model isn't the model itself, it's the ranked list of risk factors it produces. Tenure, pay tier, education, and location each point to a concrete HR intervention a manager can actually act on.