A machine-learning powered web application that classifies URLs as Benign or Phishing using only lexical and structural URL features β no DNS lookups, no page rendering, sub-5ms inference.
- Overview
- Live Demo
- Features
- Project Structure
- Dataset
- Feature Engineering
- Machine Learning Models
- Model Performance
- Dashboard Pages
- Installation
- Usage
- Reproducing Model Training
- Tech Stack
- Deployment Ideas
- License
PhishGuard AI detects phishing URLs by extracting 21 raw lexical/structural features from the URL string alone, then deriving 8 engineered features for a total of 29 features. Six scikit-learn classifiers are trained, evaluated, and compared inside an interactive Streamlit dashboard.
Why lexical-only?
- β Works offline β no network calls at inference time
- β Catches zero-day phishing before blocklists update
- β Sub-5ms prediction latency β suitable for browser extensions and DNS proxies
- β Privacy-preserving β URL is never fetched or rendered
Run locally:
streamlit run app.pyThen open http://localhost:8501 in your browser.
| Feature | Description |
|---|---|
| π Overview Dashboard | KPI cards, model comparison table, multi-dimensional radar chart |
| π Data Dashboard | Class distribution, feature correlation heatmap, feature summary stats |
| π¬ EDA & Insights | Interactive histograms, KDE plots, violin/box plots, scatter matrix, MI ranking |
| π€ Live Prediction | Manual feature sliders + sample URL picker, real-time inference across all 6 models |
| π Model Comparison | ROC curves, Precision-Recall curves, confusion matrices, CV distributions |
| π Feature Importance | Gini importance, Mutual Information scores, cumulative importance with threshold table |
| π Project Report | Full documentation: problem statement, dataset, preprocessing, engineering, results |
Phishing_URL_Detection/
β
βββ app.py # Main Streamlit application (7 pages, ~1 500 lines)
βββ train_models.py # Model training & artifact generation script
βββ requirements.txt # Python dependencies
βββ Dataset.csv # Raw URL-Phish dataset (116 600 rows)
β
βββ models/ # Pre-trained model artefacts (auto-generated)
β βββ logistic_regression.pkl
β βββ decision_tree.pkl
β βββ k_nearest_neighbours.pkl
β βββ linear_svm.pkl
β βββ gradient_boosting.pkl
β βββ random_forest.pkl
β βββ feature_cols.pkl
β βββ metrics.json
β βββ stats.json
β βββ feature_importance.json
β βββ mi_scores.json
β βββ feature_stats.json
β βββ corr_matrix.json
β βββ sample_test.csv
β
βββ data/ # (Optional) processed data directory
βββ .streamlit/ # Streamlit configuration
| Property | Value |
|---|---|
| Source | URL-Phish (synthetic, real-world feature distributions) |
| Total Records | 116,600 |
| Benign URLs | 100,000 (85.8%) |
| Phishing URLs | 16,600 (14.2%) |
| Class Ratio | ~6:1 (benign : phishing) |
| Train / Test Split | 80% / 20% stratified |
| Raw Features | 21 |
| Engineered Features | 8 |
| Total Features | 29 |
The dataset uses lexical features only β all features can be computed directly from the URL string without making any network requests.
Raw features include: url_len, dom_len, tld_len, subdom_cnt, letter_cnt, digit_cnt, special_cnt, eq_cnt, qm_cnt, amp_cnt, dot_cnt, dash_cnt, under_cnt, letter_ratio, digit_ratio, spec_ratio, is_https, slash_cnt, entropy, path_len, query_len
Eight derived features are computed on top of the raw features:
| Engineered Feature | Formula | Rationale |
|---|---|---|
complexity_index |
entropy Γ spec_ratio |
Combined entropy + special-char signal |
dom_digit_density |
digit_cnt / (dom_len + 1) |
Digit-heavy domains = auto-generated phishing |
path_url_ratio |
path_len / (url_len + 1) |
Long path in short URL β suspicious redirect |
query_url_ratio |
query_len / (url_len + 1) |
Heavy query strings β redirect obfuscation |
has_query |
int(query_len > 0) |
Binary flag for query string presence |
slash_density |
slash_cnt / (url_len + 1) |
Normalised slash count |
dot_density |
dot_cnt / (url_len + 1) |
Deeply nested sub-domains |
obfuscation_score |
eq + qm + amp + under counts |
Aggregate obfuscation character count |
3 of these engineered features rank in the top 10 by Gini importance, validating the feature engineering step.
All models are wrapped in scikit-learn Pipelines that include:
SimpleImputer(median strategy) β prevents data leakageStandardScalerβ zero-mean, unit-variance normalisation
| Model | Key Hyperparameters |
|---|---|
| Logistic Regression | C=1.0, class_weight=balanced, max_iter=1000 |
| Decision Tree | max_depth=12, class_weight=balanced, min_samples_leaf=5 |
| K-Nearest Neighbours | n_neighbors=7, n_jobs=-1 |
| Linear SVM | CalibratedClassifierCV(LinearSVC), C=0.5, class_weight=balanced |
| Gradient Boosting | n_estimators=150, max_depth=5, learning_rate=0.1, subsample=0.8 |
| Random Forest | n_estimators=200, class_weight=balanced, max_features=sqrt |
Cross-validation: 5-Fold Stratified CV is run on a 20 000-sample subset for speed.
Results are computed on the held-out 20% test set (stratified split).
| Model | Accuracy | Precision | Recall | F1 Score | ROC-AUC |
|---|---|---|---|---|---|
| π² Random Forest π | ~0.9850 | ~0.97 | ~0.94 | ~0.955 | ~0.998 |
| π Gradient Boosting | ~0.9830 | ~0.96 | ~0.93 | ~0.945 | ~0.997 |
| β‘ Linear SVM | ~0.9720 | ~0.94 | ~0.90 | ~0.920 | ~0.990 |
| π Logistic Regression | ~0.9700 | ~0.93 | ~0.89 | ~0.910 | ~0.988 |
| π³ Decision Tree | ~0.9650 | ~0.92 | ~0.88 | ~0.900 | ~0.960 |
| π K-Nearest Neighbours | ~0.9600 | ~0.90 | ~0.87 | ~0.885 | ~0.975 |
π Random Forest is the best model with the highest F1 and ROC-AUC scores.
All 6 models achieve AUC > 0.96, confirming that URL structural features are highly discriminative.
- Top-level KPI metrics (total URLs, best accuracy, F1, ROC-AUC, features, models)
- Sortable model comparison table (HTML styled)
- Multi-Dimensional Model Radar β spider chart comparing all 6 models across 5 metrics
- Class distribution donut chart
- Feature mean comparison (Benign vs Phishing) bar chart
- Interactive correlation heatmap (top 16 features)
- Feature summary statistics table
- Distributions β histogram + KDE overlay per feature, with class stats
- Violin & Box β violin plots and box plots with SD markers
- Scatter Matrix β 2D scatter with selectable X/Y features
- Class KDE β multi-feature KDE grid (up to 6 features)
- MI Ranking β Mutual Information scores bar chart, top 20
- Manual Feature Input β 21 sliders + auto-computed engineered features
- Sample URL β pick from 200 real test samples (benign or phishing)
- Primary model result card + phishing risk gauge + risk progress bar
- All-models prediction cards with per-model probability
- Consensus verdict (X/6 models agree)
- Gauge row for best model (Accuracy, Precision, Recall, F1, AUC)
- Metrics bar chart (multi-select metrics)
- ROC curves (all 6 models)
- Precision-Recall curves (all 6 models)
- Confusion matrix (model-selectable) with TP/TN/FP/FN breakdown
- 5-Fold CV F1 distribution box plots
- Random Forest Gini importance (horizontal bar + donut pie for top 10)
- Mutual Information scores (all features)
- Cumulative importance with 80%/90%/95% threshold lines and summary table
Full written documentation covering problem statement, dataset overview, preprocessing pipeline, feature engineering rationale, model results, key findings, and deployment recommendations.
- Python 3.9 or higher
- pip
# 1. Clone the repository
git clone https://github.com/SamarthGarge/Phishing_URL_Detection.git
cd Phishing_URL_Detection
# 2. (Recommended) Create a virtual environment
python -m venv venv
venv\Scripts\activate # Windows
# source venv/bin/activate # macOS / Linux
# 3. Install dependencies
pip install -r requirements.txtstreamlit>=1.32.0
pandas>=2.0.0
numpy>=1.24.0
scikit-learn>=1.3.0
plotly>=5.18.0
joblib>=1.3.0
scipy>=1.11.0
streamlit run app.pyThe app opens at http://localhost:8501.
Note: The
models/directory with pre-trained.pklfiles must be present.
If it is missing, run the training script first (see below).
Use the sidebar navigation to switch between the 7 pages:
| Sidebar Item | Page |
|---|---|
| π Overview | Model summary, radar chart |
| π Data Dashboard | Dataset visualisations |
| π¬ EDA & Insights | Exploratory analysis |
| π€ Model Prediction | Live URL prediction |
| π Model Comparison | Evaluation metrics |
| π Feature Importance | Gini & MI analysis |
| π Project Report | Full documentation |
To retrain all models from scratch:
# Make sure Dataset.csv is in the project root, then:
python train_models.pyThis script will:
- Load
Dataset.csv(or generate synthetic data if not found) - Apply IQR outlier capping on 7 continuous features
- Engineer 8 derived features
- Perform stratified 80/20 train/test split
- Train all 6 models with sklearn Pipelines
- Compute Mutual Information scores and 5-Fold CV F1
- Save all
.pklfiles and JSON artefacts tomodels/
Expected output:
π Loading dataset β¦
π Computing Mutual Information β¦
π€ Training models β¦
Model Acc Prec Rec F1 AUC CV F1
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Logistic Regression 0.9700 0.9300 0.8900 0.9100 0.9880 0.9050Β±0.003
...
π Best Model by F1: Random Forest (F1=0.9550 AUC=0.9980)
πΎ Saving artifacts β¦
β
All artifacts saved to ./models/
π Run the app with: streamlit run app.py
| Component | Technology |
|---|---|
| Web Framework | Streamlit |
| ML Library | scikit-learn |
| Data Processing | pandas, NumPy |
| Visualisation | Plotly |
| Statistical Analysis | SciPy (KDE plots) |
| Model Serialisation | joblib |
| Styling | Custom CSS (dark theme, glassmorphism, Inter/JetBrains Mono fonts) |
| Use Case | Approach |
|---|---|
| Browser Extension | Serialise RF with joblib, score URLs at navigation time (<5ms) |
| DNS Proxy Filter | Lean 5-feature model embedded in corporate DNS resolvers |
| Email Gateway | Apply model to URLs found in email bodies |
| Stack with Blocklists | ML catches zero-day phishing; PhishTank handles known-bad domains |
| Monthly Retraining | Phishing patterns evolve; fresh data prevents model drift |
| Explainability | Add SHAP values for SOC analyst trust and false positive investigation |
This project is licensed under the MIT License β see the LICENSE file for details.
Built with β€οΈ by Samarth Garge
π PhishGuard AI Β |Β Phishing URL Detection Β |Β URL-Phish Dataset (CC BY 4.0)