Solutions for the CodSoft ML internship tasks. Each task lives in its own folder with a Jupyter notebook, and saves its best model with joblib alongside a model-comparison chart.
Predict a movie's genre from its plot summary using TF-IDF features.
- Models compared: Naive Bayes, Logistic Regression, Linear SVM
- Best result: Logistic Regression — 59.1% accuracy on the official test set (27 genres)
- Dataset: Genre Classification Dataset IMDb
Detect fraudulent transactions in a heavily imbalanced dataset (~0.17% fraud), using class_weight="balanced" and F1 on the fraud class instead of raw accuracy.
- Models compared: Logistic Regression, Decision Tree, Random Forest
- Best result: Random Forest — F1 = 0.86 on the fraud class (precision 0.92, recall 0.81)
- Dataset: Credit Card Fraud Detection — too large for GitHub (144 MB), download
creditcard.csvfrom Kaggle intoTask2_CreditCardFraud/data/
Predict which bank customers will churn from demographic and account features.
- Models compared: Logistic Regression, Random Forest, Gradient Boosting
- Best result: Random Forest — F1 = 0.61 on the churn class, 84% overall accuracy
- Dataset: Churn Modelling
Classify SMS messages as spam or ham using TF-IDF features.
- Models compared: Naive Bayes, Logistic Regression, Linear SVM
- Best result: Linear SVM — F1 = 0.94 on the spam class, 98% overall accuracy
- Dataset: SMS Spam Collection
pip install pandas scikit-learn matplotlib seaborn nltk wordcloud jupyterEach notebook expects its dataset in a data/ folder next to it (included in the repo, except the Task 2 CSV noted above). Run the notebooks top to bottom.