Project Context: Employee turnover imposes substantial costs on organizations through lost expertise, productivity disruptions, and recruitment and training expenses. For this BANA 273 machine learning project, my team analyzed a publicly available Kaggle dataset of 4,653 employees across three major Indian cities to identify which employee characteristics most strongly predict voluntary churn, an employee leaving on their own, and to build a model capable of flagging at-risk employees early enough for targeted intervention.
Business Question: Which employee attributes and organizational factors most strongly influence churn, and how accurately can they be used to identify at-risk employees?
- Mid-pay tier employees churn the most, at 2.2x the odds of low-pay employees, sharpest when pay feels neither entry-level nor rewarding.
- Master's degree holders churn 2.4x more than Bachelor's holders, likely reflecting greater external mobility and a mismatch with available advancement.
- Benched employees, those without a current project assignment, churn 1.5x more than employees who have never been benched.
- Pune employees churn 1.6x more than Bangalore employees, pointing to regional labor market or workplace differences.
- Female employees show a higher churn rate than male employees, consistent across both descriptive analysis and the model.
Two models were compared on recall (the share of actual churners correctly caught), the priority metric since missing a churner costs far more than a false alarm.
The two models tell complementary stories: logistic regression surfaces linear trends and effect sizes, while the decision tree reveals how tenure, pay, and education combine into risk pathways neither variable alone would predict.
- Data prep: 4,653 records across 9 attributes from Kaggle, no missing values; categorical variables one-hot encoded for logistic regression, used directly by the decision tree; 70/30 train-test split, random state 42.
- Class imbalance: 65.6% of employees stay, only 34.4% churn, so both models used balanced class weighting and recall became the primary evaluation metric instead of accuracy.
- Modeling: logistic regression improved via GridSearchCV (L1 regularization) from 0.41 to 0.63 recall; the decision tree, pre-pruned and with its classification threshold (the probability cutoff for flagging a churner) lowered from 0.50 to 0.35, reached 0.79 recall. Both were confirmed stable via 5-fold cross-validation.
- Choosing the right evaluation metric mattered more than choosing the right model. Once recall replaced accuracy, a model that looked decent was actually catching fewer than half the churners it needed to find.
- Threshold tuning is an underused lever. Lowering the cutoff from 0.50 to 0.35 lifted the tree's recall from 0.64 to 0.79 with no additional feature engineering.
- The clearest output of a churn model isn't the model itself, it's the ranked list of risk factors it produces. Tenure, pay tier, education, and location each point to a concrete HR intervention a manager can actually act on.