A production-grade, multi-class sentiment analysis system that classifies text into Positive, Negative, and Neutral sentiments using ensemble machine learning techniques.
Built as an academic AI/ML course project demonstrating mastery of NLP fundamentals, classification algorithms, and data science best practices.
| Feature | Description |
|---|---|
| 3-Class Classification | Positive, Negative, Neutral sentiment detection |
| 6 ML Algorithms | Logistic Regression, SVM, Naive Bayes, Random Forest, XGBoost, Voting Ensemble |
| SMOTE Oversampling | Handles class imbalance via synthetic minority oversampling |
| TF-IDF Features | Unigrams + Bigrams with sublinear TF weighting |
| 10 Visualizations | Publication-quality dark-theme charts at 300 DPI |
| Real-Time Predictions | Streamlit web app with confidence scores |
| Feature Importance | Interpretable results showing top contributing words |
| Cross-Validation | 5-fold stratified CV for robust evaluation |
Sentiment_Analysis/
├── Dataset/ # Raw datasets (CSV, XLSX)
│ └── reddit_artist_posts_sentiment.csv # Primary dataset (31K+ samples)
├── src/ # Source code modules
│ ├── __init__.py
│ ├── utils.py # Configuration & utilities
│ ├── data_loading.py # Data loading & EDA
│ ├── preprocessing.py # NLP text preprocessing
│ ├── feature_engineering.py # TF-IDF feature extraction
│ ├── model_training.py # Model training with SMOTE
│ ├── evaluation.py # Metrics & model comparison
│ └── visualization.py # Professional visualizations
├── outputs/
│ ├── figures/ # Generated charts (PNG, 300 DPI)
│ └── models/ # Saved trained models (joblib)
├── main.py # Master pipeline script
├── app.py # Streamlit web app
├── requirements.txt # Python dependencies
└── README.md # This file
pip install -r requirements.txtpython main.pyThis will:
- Load and preprocess the dataset
- Train 6 ML models with SMOTE oversampling
- Evaluate all models with comprehensive metrics
- Generate 10 professional visualizations
- Save trained models for the web app
streamlit run app.py| # | Model | Key Configuration |
|---|---|---|
| 1 | Logistic Regression | C=1.0, balanced class weights, multinomial |
| 2 | Linear SVM | C=1.0, balanced weights, calibrated for probabilities |
| 3 | Multinomial Naive Bayes | Laplace smoothing (α=1.0) |
| 4 | Random Forest | 200 trees, max_depth=50, balanced weights |
| 5 | XGBoost | 200 estimators, learning_rate=0.1, max_depth=6 |
| 6 | Voting Ensemble | Soft voting of all 5 base models |
- Lowercasing — Normalize case
- URL Removal — Strip http/https links
- @Mention Removal — Remove Twitter-style handles
- Hashtag Cleaning — Keep word, remove # symbol
- HTML Entity Removal — Clean escaped characters
- Repeated Character Normalization — "sooooo" → "soo"
- Punctuation Removal — Strip special characters
- Number Removal — Remove numeric tokens
- Tokenization — Split into words
- Stop Word Removal — NLTK English + domain-specific
- Lemmatization — WordNet lemmatizer
The system generates 10 publication-quality charts:
- Class Distribution — Sentiment balance analysis
- Text Length Distribution — Word/character counts per class
- Word Clouds — Most frequent words per sentiment
- Confusion Matrices — Normalized heatmaps for all models
- Model Comparison — Grouped bar chart of all metrics
- Cross-Validation Boxplot — CV score distributions
- Feature Importance — Top 15 words per sentiment class
- ROC Curves — Multi-class OvR for all models
- Learning Curves — Train/validation accuracy vs. dataset size
- Per-Class Performance — F1-score breakdown by class
- Python 3.10+
- scikit-learn — ML models & evaluation
- XGBoost — Gradient boosting classifier
- NLTK — NLP preprocessing
- imbalanced-learn — SMOTE oversampling
- Matplotlib / Seaborn — Static visualizations
- Streamlit — Interactive web app
- Plotly — Interactive charts in web app
Academic project — MIT License