Skip to content
SAILESH4406Public

About

FraudIQ is an end-to-end, real-time banking fraud detection and investigation platform inspired by systems used internally at Visa, Mastercard, Capital One, and PayPal.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

FraudIQ — Enterprise Real-Time Fraud Intelligence & Investigation Platform

FraudIQ is an end-to-end, real-time banking fraud detection and investigation platform inspired by systems used internally at Visa, Mastercard, Capital One, and PayPal. It doesn't just score transactions — it ingests them, enriches them with behavioral context, scores them with a hybrid ML + rules risk engine, opens investigation cases, gives analysts a workspace to resolve them, and feeds every resolution back into model retraining.

Overview

Transaction
    ↓
Kafka (transactions)
    ↓
User Profile Service  →  enriches with behavioral history
    ↓
Kafka (profile_enriched_transactions)
    ↓
Fraud Engine  →  XGBoost + Isolation Forest
    ↓
Kafka (fraud_scores)
    ↓
Risk Engine  →  ML weighting + business rules + explainability
    ↓
Kafka (risk_alerts)
    ↓
Investigation Service  →  creates case, stores in MongoDB, pushes WebSocket alert
    ↓
Analyst Workspace  →  assign, note, resolve (Fraud / Legit / False Positive / Pending)
    ↓
Training Labels  →  export → retrain → compare old vs. new model

Every transaction — not just high-risk ones — becomes a reviewable case record, so analysts can override any decision, including auto-approved low-risk transactions that later turn out to be fraud.

Tech Stack

  • Backend: Python, FastAPI, JWT auth, PyMongo
  • Streaming: Apache Kafka + Zookeeper
  • Database: MongoDB
  • Machine Learning: XGBoost, Isolation Forest, scikit-learn, pandas
  • Frontend: React, Recharts, WebSockets
  • Infrastructure: Docker Compose

Project Structure

fraud_intelligence_system/
│
├── dashboard/                        # React frontend
│   └── src/
│       ├── api.js
│       ├── auth.js
│       ├── App.js / App.css / index.css
│       ├── AppLayout.js
│       ├── components/
│       │   ├── CaseDetails.js        # investigation workspace panel
│       │   ├── CaseDrawer.js
│       │   ├── CaseNotes.js
│       │   ├── CasesTable.js
│       │   ├── CasesTimelineChart.js
│       │   ├── AnalyticsCards.js
│       │   ├── LiveAlerts.js
│       │   ├── MetricsCard.js
│       │   ├── PriorityBadge.js / SeverityBadge.js
│       │   ├── ProfilesTable.js
│       │   ├── ResolutionPieChart.js / SeverityPieChart.js / RiskChart.js
│       │   ├── RiskBreakdown.js / RiskMeter.js
│       │   ├── Sidebar.js / SystemStatus.js / Timeline.js
│       │   ├── TopMerchantsTable.js / TopUsersTable.js
│       │   ├── UserProfileCard.js
│       │   └── Toast.js              # in-app toast notifications
│       └── pages/
│           ├── Dashboard.js
│           ├── CasesPage.js
│           ├── ProfilesPage.js
│           ├── AnalyticsPage.js
│           ├── ModelComparisonPage.js
│           ├── Login.js / Signup.js / Landing.js
│
├── services/
│   ├── transaction_service/
│   │   └── producer.py               # scenario-based batch/continuous generator
│   ├── user_profile_service/
│   │   ├── consumer.py / producer.py / profile_engine.py
│   ├── feature_service/
│   │   └── consumer.py
│   ├── fraud_engine/
│   │   └── consumer.py               # XGBoost + Isolation Forest scoring
│   ├── risk_service/
│   │   ├── consumer.py
│   │   ├── risk_engine.py            # ML weighting + business rules + thresholds
│   │   ├── rules.py
│   │   └── constants.py              # all thresholds/weights live here
│   └── investigation_service/
│       ├── api.py                    # FastAPI app — all REST + WebSocket endpoints
│       ├── consumer.py               # builds case records from risk_alerts
│       ├── investigation_engine.py   # case document construction, initial status
│       ├── alert_listener.py         # WebSocket broadcast for HIGH/CRITICAL alerts
│       └── websocket_manager.py
│
├── database/
│   └── mongodb/
│       └── mongo_client.py
│
├── retraining/
│   ├── export_training_data.py       # training_labels → latest_training.csv
│   ├── retrain_model.py              # trains new model, compares vs. deployed model
│   ├── latest_training.csv
│   └── model_version_history.json    # old-vs-new metrics per retraining run
│
├── models/
│   ├── xgb.pkl
│   └── iso.pkl
│
├── docker-compose.yml                # Kafka, Zookeeper, MongoDB
└── README.md

Core Features

Detection & Risk Pipeline

  • Scenario-based transaction generator (normal customer, salary day, card testing, account takeover, stolen card, money mule, crypto laundering, merchant abuse)
  • Batch mode (python producer.py --count 100) and continuous mode
  • Hybrid fraud scoring: XGBoost probability + Isolation Forest anomaly score + weighted business rules (velocity, unknown merchant/location, new device/IP, foreign transactions, impossible travel, midnight activity)
  • Every transaction — including auto-approved, low-risk ones — is now persisted as a case record, so nothing is unreviewable after the fact
  • Explainable risk breakdown attached to every case

Investigation Workflow

  • Cases page: search, filter, paginate all cases regardless of risk tier
  • Assign to analyst, add notes, resolve as Fraud / Legit / False Positive / Pending Review
  • Instant UI updates — optimistic state changes, toast notifications, no page reloads
  • Full audit log per case (who changed what, when)
  • Case locking: once resolved, only an admin can override

Analytics

  • Bounded MongoDB aggregations (no full-collection scans), short-lived server-side cache
  • Severity / priority / resolution distribution, fraud trend, daily/weekly/monthly volume
  • Top risky users and merchants

User Intelligence

  • Paginated, searchable, sortable user profiles
  • Filters: user ID, merchant, min/max average risk, minimum transaction count, fraud-history-only
  • Per-user fraud/legit case counts, average risk, last activity

Model Retraining & Comparison

  • Every case resolution becomes a training label automatically
  • Manual retraining workflow: export labels → retrain → compare → deploy
  • retrain_model.py evaluates both the previously deployed model and the newly trained one on the same validation split, so comparisons are apples-to-apples
  • Model Comparison dashboard: accuracy/precision/recall/F1 deltas, confusion matrices, retraining history

Access Control & Monitoring

  • JWT authentication, bcrypt password hashing
  • Role-based access: Admin (full platform + model monitoring + retraining) vs. Analyst (dashboard, cases, profiles, analytics)
  • System health monitoring: Model Version

Kafka Topics

Topic Published by Consumed by
transactions Transaction Producer User Profile Service
profile_enriched_transactions User Profile Service Fraud Engine
fraud_scores Fraud Engine Risk Engine
risk_alerts Risk Engine Investigation Service, Alert Listener
profile_updates Risk Engine User Profile Service (behavioral feedback)

MongoDB Collections

  • investigations — case records, risk data, resolution, audit log
  • user_profiles — behavioral aggregates per user
  • training_labels — analyst-confirmed ground truth, upserted per case
  • users — auth credentials and roles

API Reference

Auth POST /auth/signup · POST /auth/login · GET /auth/me

Cases GET /cases · GET /cases/{case_id} · PUT /cases/{case_id}/status · PUT /cases/{case_id}/assign · PUT /cases/{case_id}/notes · PUT /cases/{case_id}/resolution · GET /cases/{case_id}/explanations

Profiles GET /profiles — supports search, min_risk, max_risk, min_transactions, fraud_only, sort_by, limit, skip GET /profiles/{user_id}

Dashboard & Analytics GET /dashboard · GET /analytics · GET /metrics · GET /top-users

Model Monitoring GET /model-metrics — retraining history with old-vs-new model comparison

Realtime WS /ws — live case/alert broadcast

Setup

Prerequisites

  • Python 3.10+
  • Node.js 18+
  • Docker & Docker Compose

Infrastructure

docker compose up -d

Wait ~30 seconds after starting before launching any Python service — Kafka needs time to elect a controller and become ready; connecting too early throws KafkaTimeoutError even on a healthy setup.

Backend

pip install -r requirements.txt

# each of these runs as its own long-lived process
python -m services.transaction_service.producer --count 100
python -m services.user_profile_service.consumer
python -m services.fraud_engine.consumer
python -m services.risk_service.consumer
python -m services.investigation_service.consumer
uvicorn services.investigation_service.api:app --reload

Frontend

cd dashboard
npm install
npm start

Model Retraining

python retraining/export_training_data.py
python retraining/retrain_model.py --deploy

Results (old vs. new model metrics) appear automatically on the Model Comparison page and in /model-metrics.

Design Principles

  • Event-driven — services communicate exclusively through Kafka, no direct coupling
  • Explainable by default — every risk score ships with the factors that produced it
  • Human-in-the-loop — ML flags, analysts decide; every decision is auditable
  • Feedback-driven — resolutions become training data automatically, closing the loop between investigation and model improvement
  • Bounded by default — analytics and dashboard queries are aggregation-bound and cached, never full collection scans

Roadmap

  • Automated (scheduled) retraining
  • Model drift detection
  • Fraud network / entity-relationship graph analysis
  • Kubernetes deployment

About

FraudIQ is an end-to-end, real-time banking fraud detection and investigation platform inspired by systems used internally at Visa, Mastercard, Capital One, and PayPal.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages