Skip to content

Latest commit

Β 

History

22 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸ›‘οΈ PhishGuard API

Intelligent Phishing URL Detection β€” Powered by Machine Learning & Multi-Layered Threat Analysis

Live Demo Β· API Docs Β· Architecture


PhishGuard is a full-stack phishing URL detection platform that combines a Random Forest ML model, VirusTotal blacklist intelligence, WHOIS domain analysis, and heuristic pattern matching into a single risk score. It ships with a sleek Next.js frontend and a FastAPI backend, deployed on Vercel as a monorepo.

✨ Features

  • 4-Layer Detection Engine β€” Blacklist, Pattern Heuristics, Domain Intelligence, and ML Classifier work together for zero false-negative coverage
  • Machine Learning β€” Random Forest model trained on phishing datasets, with 17 engineered features including URL entropy, subdomain count, and brand impersonation signals
  • VirusTotal Integration β€” Real-time cross-referencing against 70+ security vendor databases
  • WHOIS Domain Age β€” Newly registered domains get escalated risk scores automatically
  • Typosquatting Detection β€” Levenshtein distance-based detection catches payytm.com, ggoogl.com, and other deceptive misspellings
  • Brand Impersonation Guard β€” Flags suspicious brand keywords on foreign domains (e.g., paypal-login.xyz)
  • Guest Mode β€” Try the API without signup; authenticated users get 100 requests/day with analytics
  • Beautiful Frontend β€” Animated UI with a WebGL splash cursor, glassmorphism cards, and real-time risk breakdowns
  • Deployed on Vercel β€” Frontend (Next.js) + Backend (Python Serverless Functions) in one repo

πŸ—οΈ Architecture

phishguard-api/
β”œβ”€β”€ frontend/                    # Monorepo root (deployed to Vercel)
β”‚   β”œβ”€β”€ api/                     # Python backend (Vercel Serverless Functions)
β”‚   β”‚   β”œβ”€β”€ index.py             # FastAPI app entry point
β”‚   β”‚   β”œβ”€β”€ requirements.txt     # Python dependencies
β”‚   β”‚   β”œβ”€β”€ phishing_model.pkl.gz  # Trained ML model (~16 MB)
β”‚   β”‚   β”œβ”€β”€ label_encoder.pkl    # Scikit-learn label encoder
β”‚   β”‚   β”œβ”€β”€ detectors/           # Detection engine modules
β”‚   β”‚   β”‚   β”œβ”€β”€ blacklist.py     # VirusTotal API integration
β”‚   β”‚   β”‚   β”œβ”€β”€ pattern.py       # URL heuristic analysis
β”‚   β”‚   β”‚   β”œβ”€β”€ domain.py        # WHOIS + SSL verification
β”‚   β”‚   β”‚   β”œβ”€β”€ ml_model.py      # Random Forest feature extraction & prediction
β”‚   β”‚   β”‚   └── scorer.py        # Multi-layer score aggregation & risk escalation
β”‚   β”‚   └── routes/              # API route handlers
β”‚   β”‚       β”œβ”€β”€ check.py         # POST /api/check-url
β”‚   β”‚       β”œβ”€β”€ auth.py          # POST /api/auth/signup, GET /api/auth/validate
β”‚   β”‚       └── analytics.py     # GET /api/analytics
β”‚   β”œβ”€β”€ src/                     # Next.js frontend
β”‚   β”‚   β”œβ”€β”€ app/
β”‚   β”‚   β”‚   β”œβ”€β”€ page.tsx         # Landing page
β”‚   β”‚   β”‚   β”œβ”€β”€ layout.tsx       # Root layout (Geist font, Vercel Analytics)
β”‚   β”‚   β”‚   β”œβ”€β”€ globals.css      # Design system & animations
β”‚   β”‚   β”‚   └── components/
β”‚   β”‚   β”‚       β”œβ”€β”€ Hero.tsx         # Animated hero section
β”‚   β”‚   β”‚       β”œβ”€β”€ URLChecker.tsx   # Interactive URL scan widget
β”‚   β”‚   β”‚       β”œβ”€β”€ HowItWorks.tsx   # Detection pipeline explainer
β”‚   β”‚   β”‚       β”œβ”€β”€ BuildStory.tsx   # Project backstory
β”‚   β”‚   β”‚       β”œβ”€β”€ APIDocs.tsx      # Interactive API documentation
β”‚   β”‚   β”‚       β”œβ”€β”€ Footer.tsx       # Footer with links
β”‚   β”‚   β”‚       └── SplashCursor.tsx # WebGL fluid cursor effect
β”‚   β”‚   └── lib/
β”‚   β”‚       └── types.ts         # TypeScript interfaces & risk level config
β”‚   β”œβ”€β”€ package.json
β”‚   β”œβ”€β”€ vercel.json              # API route rewrites
β”‚   └── next.config.ts           # Dev proxy to local FastAPI
└── .env                         # Environment variables (not committed)

πŸ”¬ Detection Pipeline

Each URL is processed through 4 independent detection layers, and their scores are aggregated with weighted fusion + risk escalation rules:

Layer Weight What It Does
Blacklist (VirusTotal) 35% Checks against 70+ AV engines. Requires β‰₯3 engines flagging at >5% ratio to avoid single-engine false positives
ML Classifier 35% Random Forest model with 17 features: URL length, entropy, digit count, brand impersonation, suspicious TLD, etc.
Pattern Heuristics 20% Regex-based checks for IP addresses, @ symbols, excessive subdomains, phishing keywords, typosquatting
Domain Intelligence 10% WHOIS domain age (<30 days = high risk), SSL certificate validation with redirect-following

Risk Escalation Rules are applied on top of the weighted score:

  • Suspicious TLD (.tk, .xyz, .click, etc.) β†’ +0.20
  • Brand impersonation on foreign domain β†’ +0.35
  • Multiple phishing keywords β†’ +0.15
  • Domain age < 30 days β†’ +0.25

Final score = max(weighted_average, strongest_signal Γ— 0.9) β€” this prevents clean layers from diluting a single high-confidence detection.

πŸš€ Getting Started

Prerequisites

1. Clone the Repository

git clone https://github.com/HariomAcharya17/phishguard-api.git
cd phishguard-api

2. Set Up Environment Variables

Create a .env file in the project root:

SUPABASE_URL=https://your-project.supabase.co
SUPABASE_KEY=your-supabase-anon-key
VIRUSTOTAL_API_KEY=your-virustotal-api-key
WHOIS_API_KEY=your-ip2whois-api-key

3. Install & Run the Backend

cd frontend/api
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn index:app --reload --port 8000

4. Install & Run the Frontend

cd frontend
npm install
npm run dev

Open http://localhost:3000 β€” the Next.js dev server proxies /api/* requests to the local FastAPI server.

5. Supabase Setup

Create two tables in your Supabase project:

api_keys

Column Type Notes
id uuid Primary key, auto-generated
email text Unique
api_key text Unique, format: pg_<uuid>
requests integer Default: 0
created_at timestamptz Default: now()

analytics

Column Type Notes
id uuid Primary key, auto-generated
api_key text Nullable (guest scans)
url text The scanned URL
risk_score float 0.0 – 1.0
is_safe boolean true if risk_score < 0.25
created_at timestamptz Default: now()

πŸ“‘ API Reference

Check URL

POST /api/check-url

Headers (optional):

x-api-key: pg_your_api_key_here

Body:

{
  "url": "https://example.com"
}

Response:

{
  "url": "https://example.com",
  "is_safe": true,
  "risk_score": 0.0512,
  "risk_level": "safe",
  "threats_detected": [],
  "ml_score": 0.0,
  "domain_age_days": 10952,
  "recommendation": "This URL appears safe. No malicious signatures were detected.",
  "breakdown": {
    "blacklist": { "score": 0.0, "threats": [], "description": "..." },
    "pattern":   { "score": 0.0, "threats": [], "description": "..." },
    "domain":    { "score": 0.0, "threats": [], "description": "..." },
    "ml":        { "score": 0.0, "threats": ["..."], "description": "..." }
  }
}

Sign Up for API Key

POST /api/auth/signup
{ "email": "you@example.com" }

Validate API Key

GET /api/auth/validate?api_key=pg_your_key

Get Analytics Dashboard

GET /api/analytics
Headers: x-api-key: pg_your_key

πŸ› οΈ Tech Stack

Component Technology
Frontend Next.js 16, React 19, TypeScript, Tailwind CSS v4, Framer Motion
Backend FastAPI, Python 3.9+, Uvicorn
ML Model Scikit-learn (Random Forest), Pandas, NumPy
Database Supabase (PostgreSQL)
Threat Intel VirusTotal API v3, IP2WHOIS API
Deployment Vercel (Serverless Functions + Edge)
Analytics Vercel Analytics

πŸ”’ Security Notes

  • API keys are never exposed to the client β€” all requests are validated server-side
  • Guest mode is available but rate-limited; authenticated users get 100 requests/day
  • The .env file containing secrets is .gitignore'd and never committed
  • CORS is configured to allow all origins (suitable for a public API)

πŸ“Š Risk Levels

Score Range Level Meaning
0.00 – 0.14 🟒 Safe No threats detected
0.15 – 0.34 🟑 Low Minor suspicious patterns
0.35 – 0.64 🟠 Medium Multiple phishing indicators
0.65 – 0.84 πŸ”΄ High Strong phishing characteristics
0.85 – 1.00 ☠️ Critical Confirmed or highly likely phishing

πŸ“ License

This project is open source. Feel free to fork, modify, and use it for your own projects.

🀝 Contributing

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

Built with ❀️ by Hariom Acharya

About

know what is fake? i will help you.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages