FraudIQ is an end-to-end, real-time banking fraud detection and investigation platform inspired by systems used internally at Visa, Mastercard, Capital One, and PayPal. It doesn't just score transactions — it ingests them, enriches them with behavioral context, scores them with a hybrid ML + rules risk engine, opens investigation cases, gives analysts a workspace to resolve them, and feeds every resolution back into model retraining.
Transaction
↓
Kafka (transactions)
↓
User Profile Service → enriches with behavioral history
↓
Kafka (profile_enriched_transactions)
↓
Fraud Engine → XGBoost + Isolation Forest
↓
Kafka (fraud_scores)
↓
Risk Engine → ML weighting + business rules + explainability
↓
Kafka (risk_alerts)
↓
Investigation Service → creates case, stores in MongoDB, pushes WebSocket alert
↓
Analyst Workspace → assign, note, resolve (Fraud / Legit / False Positive / Pending)
↓
Training Labels → export → retrain → compare old vs. new model
Every transaction — not just high-risk ones — becomes a reviewable case record, so analysts can override any decision, including auto-approved low-risk transactions that later turn out to be fraud.
- Backend: Python, FastAPI, JWT auth, PyMongo
- Streaming: Apache Kafka + Zookeeper
- Database: MongoDB
- Machine Learning: XGBoost, Isolation Forest, scikit-learn, pandas
- Frontend: React, Recharts, WebSockets
- Infrastructure: Docker Compose
fraud_intelligence_system/
│
├── dashboard/ # React frontend
│ └── src/
│ ├── api.js
│ ├── auth.js
│ ├── App.js / App.css / index.css
│ ├── AppLayout.js
│ ├── components/
│ │ ├── CaseDetails.js # investigation workspace panel
│ │ ├── CaseDrawer.js
│ │ ├── CaseNotes.js
│ │ ├── CasesTable.js
│ │ ├── CasesTimelineChart.js
│ │ ├── AnalyticsCards.js
│ │ ├── LiveAlerts.js
│ │ ├── MetricsCard.js
│ │ ├── PriorityBadge.js / SeverityBadge.js
│ │ ├── ProfilesTable.js
│ │ ├── ResolutionPieChart.js / SeverityPieChart.js / RiskChart.js
│ │ ├── RiskBreakdown.js / RiskMeter.js
│ │ ├── Sidebar.js / SystemStatus.js / Timeline.js
│ │ ├── TopMerchantsTable.js / TopUsersTable.js
│ │ ├── UserProfileCard.js
│ │ └── Toast.js # in-app toast notifications
│ └── pages/
│ ├── Dashboard.js
│ ├── CasesPage.js
│ ├── ProfilesPage.js
│ ├── AnalyticsPage.js
│ ├── ModelComparisonPage.js
│ ├── Login.js / Signup.js / Landing.js
│
├── services/
│ ├── transaction_service/
│ │ └── producer.py # scenario-based batch/continuous generator
│ ├── user_profile_service/
│ │ ├── consumer.py / producer.py / profile_engine.py
│ ├── feature_service/
│ │ └── consumer.py
│ ├── fraud_engine/
│ │ └── consumer.py # XGBoost + Isolation Forest scoring
│ ├── risk_service/
│ │ ├── consumer.py
│ │ ├── risk_engine.py # ML weighting + business rules + thresholds
│ │ ├── rules.py
│ │ └── constants.py # all thresholds/weights live here
│ └── investigation_service/
│ ├── api.py # FastAPI app — all REST + WebSocket endpoints
│ ├── consumer.py # builds case records from risk_alerts
│ ├── investigation_engine.py # case document construction, initial status
│ ├── alert_listener.py # WebSocket broadcast for HIGH/CRITICAL alerts
│ └── websocket_manager.py
│
├── database/
│ └── mongodb/
│ └── mongo_client.py
│
├── retraining/
│ ├── export_training_data.py # training_labels → latest_training.csv
│ ├── retrain_model.py # trains new model, compares vs. deployed model
│ ├── latest_training.csv
│ └── model_version_history.json # old-vs-new metrics per retraining run
│
├── models/
│ ├── xgb.pkl
│ └── iso.pkl
│
├── docker-compose.yml # Kafka, Zookeeper, MongoDB
└── README.md
- Scenario-based transaction generator (normal customer, salary day, card testing, account takeover, stolen card, money mule, crypto laundering, merchant abuse)
- Batch mode (
python producer.py --count 100) and continuous mode - Hybrid fraud scoring: XGBoost probability + Isolation Forest anomaly score + weighted business rules (velocity, unknown merchant/location, new device/IP, foreign transactions, impossible travel, midnight activity)
- Every transaction — including auto-approved, low-risk ones — is now persisted as a case record, so nothing is unreviewable after the fact
- Explainable risk breakdown attached to every case
- Cases page: search, filter, paginate all cases regardless of risk tier
- Assign to analyst, add notes, resolve as Fraud / Legit / False Positive / Pending Review
- Instant UI updates — optimistic state changes, toast notifications, no page reloads
- Full audit log per case (who changed what, when)
- Case locking: once resolved, only an admin can override
- Bounded MongoDB aggregations (no full-collection scans), short-lived server-side cache
- Severity / priority / resolution distribution, fraud trend, daily/weekly/monthly volume
- Top risky users and merchants
- Paginated, searchable, sortable user profiles
- Filters: user ID, merchant, min/max average risk, minimum transaction count, fraud-history-only
- Per-user fraud/legit case counts, average risk, last activity
- Every case resolution becomes a training label automatically
- Manual retraining workflow: export labels → retrain → compare → deploy
retrain_model.pyevaluates both the previously deployed model and the newly trained one on the same validation split, so comparisons are apples-to-apples- Model Comparison dashboard: accuracy/precision/recall/F1 deltas, confusion matrices, retraining history
- JWT authentication, bcrypt password hashing
- Role-based access: Admin (full platform + model monitoring + retraining) vs. Analyst (dashboard, cases, profiles, analytics)
- System health monitoring: Model Version
| Topic | Published by | Consumed by |
|---|---|---|
transactions |
Transaction Producer | User Profile Service |
profile_enriched_transactions |
User Profile Service | Fraud Engine |
fraud_scores |
Fraud Engine | Risk Engine |
risk_alerts |
Risk Engine | Investigation Service, Alert Listener |
profile_updates |
Risk Engine | User Profile Service (behavioral feedback) |
- investigations — case records, risk data, resolution, audit log
- user_profiles — behavioral aggregates per user
- training_labels — analyst-confirmed ground truth, upserted per case
- users — auth credentials and roles
Auth
POST /auth/signup · POST /auth/login · GET /auth/me
Cases
GET /cases · GET /cases/{case_id} · PUT /cases/{case_id}/status · PUT /cases/{case_id}/assign · PUT /cases/{case_id}/notes · PUT /cases/{case_id}/resolution · GET /cases/{case_id}/explanations
Profiles
GET /profiles — supports search, min_risk, max_risk, min_transactions, fraud_only, sort_by, limit, skip
GET /profiles/{user_id}
Dashboard & Analytics
GET /dashboard · GET /analytics · GET /metrics · GET /top-users
Model Monitoring
GET /model-metrics — retraining history with old-vs-new model comparison
Realtime
WS /ws — live case/alert broadcast
- Python 3.10+
- Node.js 18+
- Docker & Docker Compose
docker compose up -dWait ~30 seconds after starting before launching any Python service — Kafka needs time to elect a controller and become ready; connecting too early throws KafkaTimeoutError even on a healthy setup.
pip install -r requirements.txt
# each of these runs as its own long-lived process
python -m services.transaction_service.producer --count 100
python -m services.user_profile_service.consumer
python -m services.fraud_engine.consumer
python -m services.risk_service.consumer
python -m services.investigation_service.consumer
uvicorn services.investigation_service.api:app --reloadcd dashboard
npm install
npm startpython retraining/export_training_data.py
python retraining/retrain_model.py --deployResults (old vs. new model metrics) appear automatically on the Model Comparison page and in /model-metrics.
- Event-driven — services communicate exclusively through Kafka, no direct coupling
- Explainable by default — every risk score ships with the factors that produced it
- Human-in-the-loop — ML flags, analysts decide; every decision is auditable
- Feedback-driven — resolutions become training data automatically, closing the loop between investigation and model improvement
- Bounded by default — analytics and dashboard queries are aggregation-bound and cached, never full collection scans
- Automated (scheduled) retraining
- Model drift detection
- Fraud network / entity-relationship graph analysis
- Kubernetes deployment