A comprehensive Machine Learning pipeline to predict customer retention and churn.
Customer churn is a critical metric for any subscription-based business. This project provides a robust, end-to-end Machine Learning solution designed to analyze customer behavior and predict the likelihood of churn. By identifying at-risk customers, businesses can proactively implement retention strategies.
This repository fulfills Task 4 for the Data Science Internship at Sqrock IT Solutions.
👉 View the Interactive Data Report & Live Demo Here
- Data Preprocessing Pipeline: Robust handling of missing values, automated encoding of categorical features, and numerical feature scaling.
- Exploratory Data Analysis (EDA): Insightful visualizations uncovering trends in customer tenure, monthly charges, and contract types using Seaborn and Matplotlib.
- Feature Engineering: Creation of intelligent metrics such as
Average Monthly SpendandContract Type Groupingto enhance model accuracy. - Predictive Modeling: Implementation and rigorous comparison of multiple classification algorithms:
- Logistic Regression
- Decision Tree Classifier
- Random Forest Classifier (Best Performer)
- K-Nearest Neighbors (KNN)
- Model Evaluation: Thorough assessment using Accuracy, Precision, Recall, F1 Score, and Confusion Matrices.
📁 Customer_Churn_Prediction
│
├── 📓 Customer_Churn_Prediction.ipynb # Main Jupyter Notebook containing the ML pipeline
├── 📄 Telco-Customer-Churn.csv # The raw dataset
├── 🌐 index.html # Exported HTML report for GitHub Pages
└── 📝 README.md # Project documentation
Follow these steps to run the project locally on your machine.
Ensure you have Python installed, then install the required dependencies:
pip install pandas numpy matplotlib seaborn scikit-learn jupyter- Clone this repository:
git clone https://github.com/Hari2006-coder/Customer_Churn_Prediction.git
- Navigate into the directory:
cd Customer_Churn_Prediction - Launch Jupyter Notebook:
jupyter notebook
- Open
Customer_Churn_Prediction.ipynband run the cells sequentially to observe the data processing and model training phases.
After evaluating multiple models, Random Forest provided the optimal balance between Precision and Recall, making it highly effective at identifying true churn risks without excessively flagging secure customers. Detailed confusion matrices and scoring metrics are available inside the notebook.