Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ML Model Lifecycle & Drift Detection Platform

Python FastAPI MLflow Streamlit GitHub Actions Docker License

A production-grade MLOps platform that tracks experiments, detects model drift, auto-triggers retraining via CI/CD, and serves predictions through a REST API — with a live Streamlit dashboard showing model health over time.

Overview • Architecture • Features • Quick Start • API Docs • Dashboard


Overview

Most ML projects stop at the Jupyter notebook. This platform solves the hard part: what happens after deployment.

It continuously monitors a production-serving ML model for data drift and performance degradation. When drift is detected, a GitHub Actions CI/CD pipeline automatically retrains the model, logs the new experiment to MLflow, compares it against the current champion, and promotes it only if it's better.

The result is a self-healing ML system — one that doesn't silently degrade when real-world data shifts.

Problem solved: ML models degrade silently in production. Traditional CI/CD passes because the code still runs — it doesn't catch that predictions have become unreliable. This platform catches that.


Architecture

┌─────────────────────────────────────────────────────────────────────┐
│                        DATA LAYER                                   │
│   Synthetic data generator · Reference dataset · Incoming stream    │
│   Configurable drift injection for testing                          │
└──────────────────────────────┬──────────────────────────────────────┘
                               │
                               ▼
┌─────────────────────────────────────────────────────────────────────┐
│                     TRAINING PIPELINE                               │
│   Feature engineering · Model training (RandomForest + XGBoost)    │
│   MLflow experiment tracking · Model versioning · Artifact logging  │
└──────────────────────────────┬──────────────────────────────────────┘
                               │
                    ┌──────────┴──────────┐
                    ▼                     ▼
         [Champion model]        [Challenger model]
         registered in           trained on new data
         MLflow registry         compared by F1 score
                    │                     │
                    └──────────┬──────────┘
                               │ auto-promotion if challenger wins
                               ▼
┌─────────────────────────────────────────────────────────────────────┐
│                      FASTAPI SERVING LAYER                          │
│   POST /predict · GET /health · GET /model/info · GET /metrics      │
│   Prediction logging · Latency tracking · Request counting          │
└──────────────────────────────┬──────────────────────────────────────┘
                               │ (every prediction logged)
                               ▼
┌─────────────────────────────────────────────────────────────────────┐
│                     DRIFT DETECTION ENGINE                          │
│   PSI (Population Stability Index) per feature                      │
│   KS Test for distribution shift · JS Divergence                    │
│   Configurable thresholds · Drift report generation                 │
└──────────────────────────────┬──────────────────────────────────────┘
                               │
              ┌────────────────┴────────────────┐
              ▼                                 ▼
   [NO DRIFT — continue]            [DRIFT DETECTED]
                                               │
                                               ▼
┌─────────────────────────────────────────────────────────────────────┐
│                    GITHUB ACTIONS CI/CD                             │
│   Triggered by drift report · Runs retrain pipeline                 │
│   Evaluates challenger vs champion · Auto-promotes if better        │
│   Posts summary to PR / logs                                        │
└─────────────────────────────────────────────────────────────────────┘
                               │
                               ▼
┌─────────────────────────────────────────────────────────────────────┐
│                   STREAMLIT DASHBOARD                               │
│   Model health score · Drift status per feature                     │
│   MLflow experiment history · Live prediction feed                  │
│   Retraining history · Champion model metrics                       │
└─────────────────────────────────────────────────────────────────────┘

Features

MLflow Experiment Tracking

  • Every training run logged: parameters, metrics, artifacts
  • Model versioning with champion/challenger promotion
  • Automatic model comparison (F1, AUC, precision, recall)
  • Artifact storage for models, scalers, and feature configs

Drift Detection Engine

Method What it catches
PSI (Population Stability Index) Distribution shift in input features
Kolmogorov-Smirnov Test Statistical difference between reference and current data
Jensen-Shannon Divergence Symmetric measure of distribution distance
Prediction Drift Shift in model output distribution

FastAPI Serving

  • POST /predict — single and batch predictions with latency tracking
  • GET /health — liveness probe (ready for Kubernetes)
  • GET /model/info — current champion metadata from MLflow
  • GET /metrics — Prometheus-compatible metrics endpoint
  • GET /drift/status — latest drift report

GitHub Actions CI/CD

  • train.yml — manual or scheduled retraining pipeline
  • drift-check.yml — runs on push, checks for drift, triggers retrain
  • promote.yml — compares challenger vs champion, auto-promotes winner
  • Full run logs in GitHub Actions UI

Streamlit Dashboard

  • Real-time model health score (0–100)
  • Per-feature drift gauges with PSI scores
  • MLflow run history table with metric comparisons
  • Live prediction feed from the API
  • Retraining event timeline

Quick Start

Prerequisites

Python 3.10+
pip

Installation

git clone https://github.com/AwonAziz/ml-lifecycle-platform.git
cd ml-lifecycle-platform

python -m venv venv
source venv/bin/activate   # Windows: venv\Scripts\activate

pip install -r requirements.txt

Run the Full Platform

# 1. Generate training data and train the initial model
python scripts/setup.py

# 2. Start the MLflow tracking server (keep running)
mlflow ui --port 5000 &

# 3. Start the FastAPI prediction server
uvicorn src.api.server:app --reload --port 8000 &

# 4. Launch the Streamlit dashboard
streamlit run dashboard/app.py

# 5. (Optional) Inject drift and watch the system respond
python scripts/inject_drift.py

Docker

docker compose up --build

API

Predict

curl -X POST http://localhost:8000/predict \
  -H "Content-Type: application/json" \
  -d '{"features": {"feature_0": 1.2, "feature_1": -0.3, "feature_2": 0.8,
                     "feature_3": 0.1, "feature_4": -1.1}}'

Response:

{
  "prediction": 1,
  "probability": 0.847,
  "model_version": "3",
  "latency_ms": 2.3,
  "drift_status": "OK"
}

Health Check

curl http://localhost:8000/health
# {"status": "healthy", "model_loaded": true, "version": "3"}

Full interactive docs at http://localhost:8000/docs


Project Structure

ml-lifecycle-platform/
│
├── src/
│   ├── api/
│   │   ├── server.py          # FastAPI app
│   │   └── schemas.py         # Pydantic request/response models
│   ├── training/
│   │   ├── trainer.py         # MLflow-tracked training pipeline
│   │   └── features.py        # Feature engineering
│   ├── drift/
│   │   ├── detector.py        # PSI + KS + JS drift detection
│   │   └── reporter.py        # Drift report generation
│   ├── monitoring/
│   │   └── tracker.py         # Prediction logging + metrics
│   └── data/
│       └── generator.py       # Synthetic data + drift injection
│
├── dashboard/
│   └── app.py                 # Streamlit dashboard
│
├── scripts/
│   ├── setup.py               # First-run: generate data + train
│   ├── retrain.py             # Retrain + promote pipeline
│   └── inject_drift.py        # Simulate data drift for testing
│
├── .github/workflows/
│   ├── train.yml              # CI: training pipeline
│   ├── drift-check.yml        # CI: drift detection
│   └── promote.yml            # CI: champion promotion
│
├── tests/
│   ├── test_drift.py
│   ├── test_training.py
│   └── test_api.py
│
├── config/settings.py
├── requirements.txt
├── Dockerfile
└── docker-compose.yml

Results

Metric Value
API prediction latency < 5ms (p99)
Drift detection sensitivity PSI > 0.1 triggers alert
Model promotion threshold Challenger F1 > Champion F1
CI/CD pipeline duration ~2 minutes end-to-end
Dashboard refresh rate Every 5 seconds

Tech Stack

Layer Technology
Model Training Scikit-Learn, XGBoost, MLflow
API Serving FastAPI, Uvicorn, Pydantic
Drift Detection SciPy, NumPy (PSI + KS + JS)
Dashboard Streamlit, Plotly
CI/CD GitHub Actions
Containerization Docker, Docker Compose
Experiment Tracking MLflow

License

MIT — see LICENSE


Built by Awon Aziz · ML & AIOps Engineer

About

A FastAPI-served ML model with MLflow experiment tracking, automated data drift detection (using Evidently AI), GitHub Actions CI/CD that auto-retriggers retraining when drift is detected, and a Streamlit dashboard showing model health over time.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages