Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🎭 Multi-Class Sentiment Analysis System

A production-grade, multi-class sentiment analysis system that classifies text into Positive, Negative, and Neutral sentiments using ensemble machine learning techniques.

Built as an academic AI/ML course project demonstrating mastery of NLP fundamentals, classification algorithms, and data science best practices.


🎯 Project Highlights

Feature Description
3-Class Classification Positive, Negative, Neutral sentiment detection
6 ML Algorithms Logistic Regression, SVM, Naive Bayes, Random Forest, XGBoost, Voting Ensemble
SMOTE Oversampling Handles class imbalance via synthetic minority oversampling
TF-IDF Features Unigrams + Bigrams with sublinear TF weighting
10 Visualizations Publication-quality dark-theme charts at 300 DPI
Real-Time Predictions Streamlit web app with confidence scores
Feature Importance Interpretable results showing top contributing words
Cross-Validation 5-fold stratified CV for robust evaluation

📂 Project Structure

Sentiment_Analysis/
├── Dataset/                          # Raw datasets (CSV, XLSX)
│   └── reddit_artist_posts_sentiment.csv   # Primary dataset (31K+ samples)
├── src/                              # Source code modules
│   ├── __init__.py
│   ├── utils.py                      # Configuration & utilities
│   ├── data_loading.py               # Data loading & EDA
│   ├── preprocessing.py              # NLP text preprocessing
│   ├── feature_engineering.py        # TF-IDF feature extraction
│   ├── model_training.py             # Model training with SMOTE
│   ├── evaluation.py                 # Metrics & model comparison
│   └── visualization.py             # Professional visualizations
├── outputs/
│   ├── figures/                      # Generated charts (PNG, 300 DPI)
│   └── models/                       # Saved trained models (joblib)
├── main.py                           # Master pipeline script
├── app.py                            # Streamlit web app
├── requirements.txt                  # Python dependencies
└── README.md                         # This file

🚀 Quick Start

1. Install Dependencies

pip install -r requirements.txt

2. Run the Full Pipeline

python main.py

This will:

  • Load and preprocess the dataset
  • Train 6 ML models with SMOTE oversampling
  • Evaluate all models with comprehensive metrics
  • Generate 10 professional visualizations
  • Save trained models for the web app

3. Launch the Web App

streamlit run app.py

📊 Models Implemented

# Model Key Configuration
1 Logistic Regression C=1.0, balanced class weights, multinomial
2 Linear SVM C=1.0, balanced weights, calibrated for probabilities
3 Multinomial Naive Bayes Laplace smoothing (α=1.0)
4 Random Forest 200 trees, max_depth=50, balanced weights
5 XGBoost 200 estimators, learning_rate=0.1, max_depth=6
6 Voting Ensemble Soft voting of all 5 base models

🔧 NLP Pipeline

  1. Lowercasing — Normalize case
  2. URL Removal — Strip http/https links
  3. @Mention Removal — Remove Twitter-style handles
  4. Hashtag Cleaning — Keep word, remove # symbol
  5. HTML Entity Removal — Clean escaped characters
  6. Repeated Character Normalization — "sooooo" → "soo"
  7. Punctuation Removal — Strip special characters
  8. Number Removal — Remove numeric tokens
  9. Tokenization — Split into words
  10. Stop Word Removal — NLTK English + domain-specific
  11. Lemmatization — WordNet lemmatizer

📈 Visualizations

The system generates 10 publication-quality charts:

  1. Class Distribution — Sentiment balance analysis
  2. Text Length Distribution — Word/character counts per class
  3. Word Clouds — Most frequent words per sentiment
  4. Confusion Matrices — Normalized heatmaps for all models
  5. Model Comparison — Grouped bar chart of all metrics
  6. Cross-Validation Boxplot — CV score distributions
  7. Feature Importance — Top 15 words per sentiment class
  8. ROC Curves — Multi-class OvR for all models
  9. Learning Curves — Train/validation accuracy vs. dataset size
  10. Per-Class Performance — F1-score breakdown by class

🛠 Technology Stack

  • Python 3.10+
  • scikit-learn — ML models & evaluation
  • XGBoost — Gradient boosting classifier
  • NLTK — NLP preprocessing
  • imbalanced-learn — SMOTE oversampling
  • Matplotlib / Seaborn — Static visualizations
  • Streamlit — Interactive web app
  • Plotly — Interactive charts in web app

📝 License

Academic project — MIT License

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages