An end-to-end, production-ready Machine Learning system designed to predict customer churn in subscription-based businesses such as Telecom, SaaS, and Banking.
Live Interactive Web Dashboard deployed on Streamlit Cloud
🌐 Live Web App: https://churnmodel-by-vishal-dubey.streamlit.app/
Click the link above to access the live app, adjust customer profile parameters, predict churn risk probabilities in real-time, and view tailored business retention recommendations.
Customer churn directly impacts business revenue. Failing to identify churn-prone customers leads to significant financial loss, while timely detection enables proactive retention strategies.
Predict whether a customer will churn (1) or stay (0), enabling businesses to take data-driven retention actions before losing valuable accounts.
- ✅ End-to-End ML Pipeline: Complete flow from raw data EDA to interactive web deployment.
- ✅ Data Leakage Prevention: Removed post-churn features like total charges and location noise to ensure high real-world accuracy.
- ✅ Class Imbalance Management: Handled ~26% baseline churn distribution effectively.
- ✅ Recall-Optimized: Prioritized high recall for churners to minimize costly false negatives.
- ✅ Ensemble Learning: Utilized a Soft Voting Classifier combining Logistic Regression, Random Forest, and Gradient Boosting.
- ✅ Interactive Dashboard: Built with Streamlit offering real-time prediction sliders, risk badges, and business suggestions.
The project follows a structured, industry-aligned ML workflow:
- Exploratory Data Analysis (EDA): Deep dive into customer demographics, contract types, and billing behaviors.
- Hypothesis Testing: Validated core assumptions regarding contract duration, monthly charges, and churn rate.
- Data Leakage Prevention: Filtered out collinear features (
TotalCharges,TotalRevenue) available only after customer lifecycle ends. - Feature Engineering & Encoding: Applied categorical target encoding and standard scaling.
- Baseline Modeling: Trained Logistic Regression as a baseline model.
- Advanced & Ensemble Modeling: Built Random Forest, Gradient Boosting, and a Soft Voting Classifier.
- Business Metrics Optimization: Tuned decision thresholds focusing on recall over raw accuracy.
- Deployment: Packaged and deployed via Streamlit Cloud for instant accessibility.
- Dataset Size: ~7,000+ customer records
- Domain: Telecom customer behavior
ChurnLabel1→ Customer Churned ❌0→ Customer Retained ✅
| Category | Key Features |
|---|---|
| 👤 Demographics | Age, Gender, Senior Citizen Status, Dependents |
| 📊 Account & Contract | Contract Type (Month-to-month, One year, Two year), Tenure (Months) |
| 💰 Usage & Billing | Monthly Charges, Payment Method, Internet Service Type |
| ⭐ Customer Value | CLTV (Customer Lifetime Value), Satisfaction Score |
- 🔹 Contract Sensitivity: Customers on Month-to-month contracts exhibit the highest churn probability.
- 🔹 Tenure Impact: Low-tenure customers (< 12 months) are significantly more prone to churn.
- 🔹 Pricing Pressure: Higher monthly charges positively correlate with elevated churn risk.
- 🔹 Satisfaction Score: Low customer satisfaction scores serve as the strongest early warning indicator.
| Model | Accuracy | Recall (Churn) | ROC-AUC | Status |
|---|---|---|---|---|
| Logistic Regression | ~88% | Good | ~0.96 | Baseline |
| Random Forest | ~88% | Moderate | ~0.95 | Tree-based |
| Gradient Boosting | ~95% | High | ~0.99 | Strong Performer |
| Voting Classifier (Ensemble) | ~95% | Highest 🔥 | ~0.98 | Production Model |
In customer churn prediction, a false negative (missing a churner) is far more expensive than a false positive (sending a discount to a loyal customer). Therefore, our optimization strategy explicitly maximizes Recall for Churned Customers, ensuring retention teams can proactively intervene.
Churn_Model/
│
├── assets/
│ └── app_preview.png # Web Application Screenshot Showcase
├── data/ # Dataset files
├── notebooks/ # EDA & Model Training Jupyter Notebooks
├── model/
│ ├── churn_model.pkl # Trained Soft Voting Ensemble Model
│ └── feature_names.pkl # Feature Metadata
├── app.py # Main Streamlit Web Application
├── requirements.txt # Python Dependencies
└── README.md # Documentation
git clone https://github.com/Vishaldubey2210/Churn_Model.git
cd Churn_Modelpip install -r requirements.txtstreamlit run app.pyThe application will launch locally at http://localhost:8501.
- Integrate SHAP (SHapley Additive exPlanations) for local model interpretability.
- Automated Hyperparameter Optimization using Optuna.
- REST API development via FastAPI.
- CI/CD pipeline integration with GitHub Actions.
Vishal Kumar
Aspiring Machine Learning Engineer
- 🌐 Live Application: https://churnmodel-by-vishal-dubey.streamlit.app/
- 💻 GitHub: @Vishaldubey2210