Skip to content

Latest commit

Β 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸ”’ PhishGuard AI β€” Phishing URL Detection System

Python Streamlit scikit-learn Plotly SciPy License: MIT

A machine-learning powered web application that classifies URLs as Benign or Phishing using only lexical and structural URL features β€” no DNS lookups, no page rendering, sub-5ms inference.


πŸ“‹ Table of Contents


🎯 Overview

PhishGuard AI detects phishing URLs by extracting 21 raw lexical/structural features from the URL string alone, then deriving 8 engineered features for a total of 29 features. Six scikit-learn classifiers are trained, evaluated, and compared inside an interactive Streamlit dashboard.

Why lexical-only?

  • βœ… Works offline β€” no network calls at inference time
  • βœ… Catches zero-day phishing before blocklists update
  • βœ… Sub-5ms prediction latency β€” suitable for browser extensions and DNS proxies
  • βœ… Privacy-preserving β€” URL is never fetched or rendered

πŸš€ Live Demo

Run locally:

streamlit run app.py

Then open http://localhost:8501 in your browser.


✨ Features

Feature Description
🏠 Overview Dashboard KPI cards, model comparison table, multi-dimensional radar chart
πŸ“Š Data Dashboard Class distribution, feature correlation heatmap, feature summary stats
πŸ”¬ EDA & Insights Interactive histograms, KDE plots, violin/box plots, scatter matrix, MI ranking
πŸ€– Live Prediction Manual feature sliders + sample URL picker, real-time inference across all 6 models
πŸ“ˆ Model Comparison ROC curves, Precision-Recall curves, confusion matrices, CV distributions
πŸ” Feature Importance Gini importance, Mutual Information scores, cumulative importance with threshold table
πŸ“‹ Project Report Full documentation: problem statement, dataset, preprocessing, engineering, results

πŸ“ Project Structure

Phishing_URL_Detection/
β”‚
β”œβ”€β”€ app.py                  # Main Streamlit application (7 pages, ~1 500 lines)
β”œβ”€β”€ train_models.py         # Model training & artifact generation script
β”œβ”€β”€ requirements.txt        # Python dependencies
β”œβ”€β”€ Dataset.csv             # Raw URL-Phish dataset (116 600 rows)
β”‚
β”œβ”€β”€ models/                 # Pre-trained model artefacts (auto-generated)
β”‚   β”œβ”€β”€ logistic_regression.pkl
β”‚   β”œβ”€β”€ decision_tree.pkl
β”‚   β”œβ”€β”€ k_nearest_neighbours.pkl
β”‚   β”œβ”€β”€ linear_svm.pkl
β”‚   β”œβ”€β”€ gradient_boosting.pkl
β”‚   β”œβ”€β”€ random_forest.pkl
β”‚   β”œβ”€β”€ feature_cols.pkl
β”‚   β”œβ”€β”€ metrics.json
β”‚   β”œβ”€β”€ stats.json
β”‚   β”œβ”€β”€ feature_importance.json
β”‚   β”œβ”€β”€ mi_scores.json
β”‚   β”œβ”€β”€ feature_stats.json
β”‚   β”œβ”€β”€ corr_matrix.json
β”‚   └── sample_test.csv
β”‚
β”œβ”€β”€ data/                   # (Optional) processed data directory
└── .streamlit/             # Streamlit configuration

πŸ“‚ Dataset

Property Value
Source URL-Phish (synthetic, real-world feature distributions)
Total Records 116,600
Benign URLs 100,000 (85.8%)
Phishing URLs 16,600 (14.2%)
Class Ratio ~6:1 (benign : phishing)
Train / Test Split 80% / 20% stratified
Raw Features 21
Engineered Features 8
Total Features 29

The dataset uses lexical features only β€” all features can be computed directly from the URL string without making any network requests.

Raw features include: url_len, dom_len, tld_len, subdom_cnt, letter_cnt, digit_cnt, special_cnt, eq_cnt, qm_cnt, amp_cnt, dot_cnt, dash_cnt, under_cnt, letter_ratio, digit_ratio, spec_ratio, is_https, slash_cnt, entropy, path_len, query_len


πŸ”§ Feature Engineering

Eight derived features are computed on top of the raw features:

Engineered Feature Formula Rationale
complexity_index entropy Γ— spec_ratio Combined entropy + special-char signal
dom_digit_density digit_cnt / (dom_len + 1) Digit-heavy domains = auto-generated phishing
path_url_ratio path_len / (url_len + 1) Long path in short URL β†’ suspicious redirect
query_url_ratio query_len / (url_len + 1) Heavy query strings β†’ redirect obfuscation
has_query int(query_len > 0) Binary flag for query string presence
slash_density slash_cnt / (url_len + 1) Normalised slash count
dot_density dot_cnt / (url_len + 1) Deeply nested sub-domains
obfuscation_score eq + qm + amp + under counts Aggregate obfuscation character count

3 of these engineered features rank in the top 10 by Gini importance, validating the feature engineering step.


πŸ€– Machine Learning Models

All models are wrapped in scikit-learn Pipelines that include:

  1. SimpleImputer (median strategy) β€” prevents data leakage
  2. StandardScaler β€” zero-mean, unit-variance normalisation
Model Key Hyperparameters
Logistic Regression C=1.0, class_weight=balanced, max_iter=1000
Decision Tree max_depth=12, class_weight=balanced, min_samples_leaf=5
K-Nearest Neighbours n_neighbors=7, n_jobs=-1
Linear SVM CalibratedClassifierCV(LinearSVC), C=0.5, class_weight=balanced
Gradient Boosting n_estimators=150, max_depth=5, learning_rate=0.1, subsample=0.8
Random Forest n_estimators=200, class_weight=balanced, max_features=sqrt

Cross-validation: 5-Fold Stratified CV is run on a 20 000-sample subset for speed.


πŸ“Š Model Performance

Results are computed on the held-out 20% test set (stratified split).

Model Accuracy Precision Recall F1 Score ROC-AUC
🌲 Random Forest πŸ† ~0.9850 ~0.97 ~0.94 ~0.955 ~0.998
πŸš€ Gradient Boosting ~0.9830 ~0.96 ~0.93 ~0.945 ~0.997
⚑ Linear SVM ~0.9720 ~0.94 ~0.90 ~0.920 ~0.990
πŸ“ˆ Logistic Regression ~0.9700 ~0.93 ~0.89 ~0.910 ~0.988
🌳 Decision Tree ~0.9650 ~0.92 ~0.88 ~0.900 ~0.960
πŸ” K-Nearest Neighbours ~0.9600 ~0.90 ~0.87 ~0.885 ~0.975

πŸ† Random Forest is the best model with the highest F1 and ROC-AUC scores.
All 6 models achieve AUC > 0.96, confirming that URL structural features are highly discriminative.


πŸ–₯️ Dashboard Pages

🏠 Overview

  • Top-level KPI metrics (total URLs, best accuracy, F1, ROC-AUC, features, models)
  • Sortable model comparison table (HTML styled)
  • Multi-Dimensional Model Radar β€” spider chart comparing all 6 models across 5 metrics

πŸ“Š Data Dashboard

  • Class distribution donut chart
  • Feature mean comparison (Benign vs Phishing) bar chart
  • Interactive correlation heatmap (top 16 features)
  • Feature summary statistics table

πŸ”¬ EDA & Insights (5 tabs)

  • Distributions β€” histogram + KDE overlay per feature, with class stats
  • Violin & Box β€” violin plots and box plots with SD markers
  • Scatter Matrix β€” 2D scatter with selectable X/Y features
  • Class KDE β€” multi-feature KDE grid (up to 6 features)
  • MI Ranking β€” Mutual Information scores bar chart, top 20

πŸ€– Model Prediction

  • Manual Feature Input β€” 21 sliders + auto-computed engineered features
  • Sample URL β€” pick from 200 real test samples (benign or phishing)
  • Primary model result card + phishing risk gauge + risk progress bar
  • All-models prediction cards with per-model probability
  • Consensus verdict (X/6 models agree)

πŸ“ˆ Model Comparison (5 tabs)

  • Gauge row for best model (Accuracy, Precision, Recall, F1, AUC)
  • Metrics bar chart (multi-select metrics)
  • ROC curves (all 6 models)
  • Precision-Recall curves (all 6 models)
  • Confusion matrix (model-selectable) with TP/TN/FP/FN breakdown
  • 5-Fold CV F1 distribution box plots

πŸ” Feature Importance (3 tabs)

  • Random Forest Gini importance (horizontal bar + donut pie for top 10)
  • Mutual Information scores (all features)
  • Cumulative importance with 80%/90%/95% threshold lines and summary table

πŸ“‹ Project Report

Full written documentation covering problem statement, dataset overview, preprocessing pipeline, feature engineering rationale, model results, key findings, and deployment recommendations.


βš™οΈ Installation

Prerequisites

  • Python 3.9 or higher
  • pip

Steps

# 1. Clone the repository
git clone https://github.com/SamarthGarge/Phishing_URL_Detection.git
cd Phishing_URL_Detection

# 2. (Recommended) Create a virtual environment
python -m venv venv
venv\Scripts\activate        # Windows
# source venv/bin/activate   # macOS / Linux

# 3. Install dependencies
pip install -r requirements.txt

Dependencies

streamlit>=1.32.0
pandas>=2.0.0
numpy>=1.24.0
scikit-learn>=1.3.0
plotly>=5.18.0
joblib>=1.3.0
scipy>=1.11.0

πŸš€ Usage

Run the Streamlit App

streamlit run app.py

The app opens at http://localhost:8501.

Note: The models/ directory with pre-trained .pkl files must be present.
If it is missing, run the training script first (see below).

Navigate the App

Use the sidebar navigation to switch between the 7 pages:

Sidebar Item Page
🏠 Overview Model summary, radar chart
πŸ“Š Data Dashboard Dataset visualisations
πŸ”¬ EDA & Insights Exploratory analysis
πŸ€– Model Prediction Live URL prediction
πŸ“ˆ Model Comparison Evaluation metrics
πŸ” Feature Importance Gini & MI analysis
πŸ“‹ Project Report Full documentation

πŸ‹οΈ Reproducing Model Training

To retrain all models from scratch:

# Make sure Dataset.csv is in the project root, then:
python train_models.py

This script will:

  1. Load Dataset.csv (or generate synthetic data if not found)
  2. Apply IQR outlier capping on 7 continuous features
  3. Engineer 8 derived features
  4. Perform stratified 80/20 train/test split
  5. Train all 6 models with sklearn Pipelines
  6. Compute Mutual Information scores and 5-Fold CV F1
  7. Save all .pkl files and JSON artefacts to models/

Expected output:

πŸ“‚ Loading dataset …
πŸ“ Computing Mutual Information …
πŸ€– Training models …
Model                     Acc    Prec     Rec      F1     AUC      CV F1
────────────────────────────────────────────────────────────────────────
Logistic Regression     0.9700  0.9300  0.8900  0.9100  0.9880  0.9050Β±0.003
...
πŸ† Best Model by F1: Random Forest  (F1=0.9550  AUC=0.9980)
πŸ’Ύ Saving artifacts …
   βœ…  All artifacts saved to ./models/
πŸš€ Run the app with:  streamlit run app.py

πŸ› οΈ Tech Stack

Component Technology
Web Framework Streamlit
ML Library scikit-learn
Data Processing pandas, NumPy
Visualisation Plotly
Statistical Analysis SciPy (KDE plots)
Model Serialisation joblib
Styling Custom CSS (dark theme, glassmorphism, Inter/JetBrains Mono fonts)

🚒 Deployment Ideas

Use Case Approach
Browser Extension Serialise RF with joblib, score URLs at navigation time (<5ms)
DNS Proxy Filter Lean 5-feature model embedded in corporate DNS resolvers
Email Gateway Apply model to URLs found in email bodies
Stack with Blocklists ML catches zero-day phishing; PhishTank handles known-bad domains
Monthly Retraining Phishing patterns evolve; fresh data prevents model drift
Explainability Add SHAP values for SOC analyst trust and false positive investigation

πŸ“„ License

This project is licensed under the MIT License β€” see the LICENSE file for details.


Built with ❀️ by Samarth Garge

πŸ”’ PhishGuard AI Β |Β  Phishing URL Detection Β |Β  URL-Phish Dataset (CC BY 4.0)

About

Phishing URL Detection is a Python project that helps identify and classify phishing websites by analyzing URLs. It uses machine learning to detect malicious links, making it easier to protect users from online scams. This project is useful for anyone interested in cybersecurity or automated threat detection.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages