Intelligent Phishing URL Detection β Powered by Machine Learning & Multi-Layered Threat Analysis
Live Demo Β· API Docs Β· Architecture
PhishGuard is a full-stack phishing URL detection platform that combines a Random Forest ML model, VirusTotal blacklist intelligence, WHOIS domain analysis, and heuristic pattern matching into a single risk score. It ships with a sleek Next.js frontend and a FastAPI backend, deployed on Vercel as a monorepo.
- 4-Layer Detection Engine β Blacklist, Pattern Heuristics, Domain Intelligence, and ML Classifier work together for zero false-negative coverage
- Machine Learning β Random Forest model trained on phishing datasets, with 17 engineered features including URL entropy, subdomain count, and brand impersonation signals
- VirusTotal Integration β Real-time cross-referencing against 70+ security vendor databases
- WHOIS Domain Age β Newly registered domains get escalated risk scores automatically
- Typosquatting Detection β Levenshtein distance-based detection catches
payytm.com,ggoogl.com, and other deceptive misspellings - Brand Impersonation Guard β Flags suspicious brand keywords on foreign domains (e.g.,
paypal-login.xyz) - Guest Mode β Try the API without signup; authenticated users get 100 requests/day with analytics
- Beautiful Frontend β Animated UI with a WebGL splash cursor, glassmorphism cards, and real-time risk breakdowns
- Deployed on Vercel β Frontend (Next.js) + Backend (Python Serverless Functions) in one repo
phishguard-api/
βββ frontend/ # Monorepo root (deployed to Vercel)
β βββ api/ # Python backend (Vercel Serverless Functions)
β β βββ index.py # FastAPI app entry point
β β βββ requirements.txt # Python dependencies
β β βββ phishing_model.pkl.gz # Trained ML model (~16 MB)
β β βββ label_encoder.pkl # Scikit-learn label encoder
β β βββ detectors/ # Detection engine modules
β β β βββ blacklist.py # VirusTotal API integration
β β β βββ pattern.py # URL heuristic analysis
β β β βββ domain.py # WHOIS + SSL verification
β β β βββ ml_model.py # Random Forest feature extraction & prediction
β β β βββ scorer.py # Multi-layer score aggregation & risk escalation
β β βββ routes/ # API route handlers
β β βββ check.py # POST /api/check-url
β β βββ auth.py # POST /api/auth/signup, GET /api/auth/validate
β β βββ analytics.py # GET /api/analytics
β βββ src/ # Next.js frontend
β β βββ app/
β β β βββ page.tsx # Landing page
β β β βββ layout.tsx # Root layout (Geist font, Vercel Analytics)
β β β βββ globals.css # Design system & animations
β β β βββ components/
β β β βββ Hero.tsx # Animated hero section
β β β βββ URLChecker.tsx # Interactive URL scan widget
β β β βββ HowItWorks.tsx # Detection pipeline explainer
β β β βββ BuildStory.tsx # Project backstory
β β β βββ APIDocs.tsx # Interactive API documentation
β β β βββ Footer.tsx # Footer with links
β β β βββ SplashCursor.tsx # WebGL fluid cursor effect
β β βββ lib/
β β βββ types.ts # TypeScript interfaces & risk level config
β βββ package.json
β βββ vercel.json # API route rewrites
β βββ next.config.ts # Dev proxy to local FastAPI
βββ .env # Environment variables (not committed)
Each URL is processed through 4 independent detection layers, and their scores are aggregated with weighted fusion + risk escalation rules:
| Layer | Weight | What It Does |
|---|---|---|
| Blacklist (VirusTotal) | 35% | Checks against 70+ AV engines. Requires β₯3 engines flagging at >5% ratio to avoid single-engine false positives |
| ML Classifier | 35% | Random Forest model with 17 features: URL length, entropy, digit count, brand impersonation, suspicious TLD, etc. |
| Pattern Heuristics | 20% | Regex-based checks for IP addresses, @ symbols, excessive subdomains, phishing keywords, typosquatting |
| Domain Intelligence | 10% | WHOIS domain age (<30 days = high risk), SSL certificate validation with redirect-following |
Risk Escalation Rules are applied on top of the weighted score:
- Suspicious TLD (
.tk,.xyz,.click, etc.) β +0.20 - Brand impersonation on foreign domain β +0.35
- Multiple phishing keywords β +0.15
- Domain age < 30 days β +0.25
Final score = max(weighted_average, strongest_signal Γ 0.9) β this prevents clean layers from diluting a single high-confidence detection.
- Node.js β₯ 18
- Python β₯ 3.9
- API keys for VirusTotal, IP2WHOIS, and a Supabase project
git clone https://github.com/HariomAcharya17/phishguard-api.git
cd phishguard-apiCreate a .env file in the project root:
SUPABASE_URL=https://your-project.supabase.co
SUPABASE_KEY=your-supabase-anon-key
VIRUSTOTAL_API_KEY=your-virustotal-api-key
WHOIS_API_KEY=your-ip2whois-api-keycd frontend/api
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn index:app --reload --port 8000cd frontend
npm install
npm run devOpen http://localhost:3000 β the Next.js dev server proxies /api/* requests to the local FastAPI server.
Create two tables in your Supabase project:
api_keys
| Column | Type | Notes |
|---|---|---|
| id | uuid | Primary key, auto-generated |
| text | Unique | |
| api_key | text | Unique, format: pg_<uuid> |
| requests | integer | Default: 0 |
| created_at | timestamptz | Default: now() |
analytics
| Column | Type | Notes |
|---|---|---|
| id | uuid | Primary key, auto-generated |
| api_key | text | Nullable (guest scans) |
| url | text | The scanned URL |
| risk_score | float | 0.0 β 1.0 |
| is_safe | boolean | true if risk_score < 0.25 |
| created_at | timestamptz | Default: now() |
POST /api/check-url
Headers (optional):
x-api-key: pg_your_api_key_here
Body:
{
"url": "https://example.com"
}Response:
{
"url": "https://example.com",
"is_safe": true,
"risk_score": 0.0512,
"risk_level": "safe",
"threats_detected": [],
"ml_score": 0.0,
"domain_age_days": 10952,
"recommendation": "This URL appears safe. No malicious signatures were detected.",
"breakdown": {
"blacklist": { "score": 0.0, "threats": [], "description": "..." },
"pattern": { "score": 0.0, "threats": [], "description": "..." },
"domain": { "score": 0.0, "threats": [], "description": "..." },
"ml": { "score": 0.0, "threats": ["..."], "description": "..." }
}
}POST /api/auth/signup
{ "email": "you@example.com" }GET /api/auth/validate?api_key=pg_your_key
GET /api/analytics
Headers: x-api-key: pg_your_key
| Component | Technology |
|---|---|
| Frontend | Next.js 16, React 19, TypeScript, Tailwind CSS v4, Framer Motion |
| Backend | FastAPI, Python 3.9+, Uvicorn |
| ML Model | Scikit-learn (Random Forest), Pandas, NumPy |
| Database | Supabase (PostgreSQL) |
| Threat Intel | VirusTotal API v3, IP2WHOIS API |
| Deployment | Vercel (Serverless Functions + Edge) |
| Analytics | Vercel Analytics |
- API keys are never exposed to the client β all requests are validated server-side
- Guest mode is available but rate-limited; authenticated users get 100 requests/day
- The
.envfile containing secrets is.gitignore'd and never committed - CORS is configured to allow all origins (suitable for a public API)
| Score Range | Level | Meaning |
|---|---|---|
| 0.00 β 0.14 | π’ Safe | No threats detected |
| 0.15 β 0.34 | π‘ Low | Minor suspicious patterns |
| 0.35 β 0.64 | π Medium | Multiple phishing indicators |
| 0.65 β 0.84 | π΄ High | Strong phishing characteristics |
| 0.85 β 1.00 | β οΈ Critical | Confirmed or highly likely phishing |
This project is open source. Feel free to fork, modify, and use it for your own projects.
- Fork the repository
- Create your feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
Built with β€οΈ by Hariom Acharya