Champion-challenger ML deployment with traffic splitting, A/B testing using alpha-spending (O'Brien-Fleming), and automated rollback.
flowchart TB
subgraph Router["Traffic Router"]
REQ[Request] --> SPLIT{Traffic Splitter}
SPLIT -->|90%| CHAMP[Champion Model]
SPLIT -->|10%| CHALL[Challenger Model]
end
subgraph Testing["A/B Testing"]
CHAMP --> METRICS[Metrics Collector]
CHALL --> METRICS
METRICS --> SEQ[Sequential Test]
SEQ -->|O'Brien-Fleming| DECISION{Decision}
end
subgraph Rollback["Auto Rollback"]
DECISION -->|Degradation| ROLL[Rollback Trigger]
GUARD[Guardrail Check] --> ROLL
ROLL --> SPLIT
end
DECISION -->|Champion Wins| KEEP[Keep Champion]
DECISION -->|Challenger Wins| ADVANCE[Advance Stage]
| Stage | Champion | Challenger | Purpose |
|---|---|---|---|
| Initial (Canary) | 90% | 10% | Detect major issues |
| Validation | 50% | 50% | Collect A/B data |
| Full Rollout | 0% | 100% | Complete migration |
Uses O'Brien-Fleming spending function for valid early stopping:
Boundary(t) = z_alpha / sqrt(t)
where t = information fraction (samples collected / total planned)
Benefits:
- Check results multiple times without inflating false positive rate
- Stop early if challenger is clearly better or worse
- Conservative early (high bar), less conservative later
Triggers on:
- Guardrail violations: Latency > 100ms, error rate > 1%
- Metric degradation: Primary metric drops > 5%
- Statistical degradation: Z-statistic exceeds negative boundary
# Install
make install
# Run tests
make test
# Start server
make serve| Endpoint | Method | Description |
|---|---|---|
/predict |
POST | Route prediction to champion/challenger |
/record-metrics |
POST | Record outcome metrics |
/traffic-split |
GET/POST | View/update traffic split |
/advance-stage |
POST | Move to next deployment stage |
/rollback |
POST | Trigger manual rollback |
/test-result |
GET | Get A/B test analysis |
/metrics-summary |
GET | Get metrics by variant |
curl -X POST http://localhost:8000/predict \
-H "Content-Type: application/json" \
-d '{"features": {"amount": 100}, "session_id": "user_123"}'Response:
{
"prediction": "positive",
"probability": 0.82,
"variant": "champion",
"request_id": "abc-123"
}curl -X POST http://localhost:8000/record-metrics \
-H "Content-Type: application/json" \
-d '{
"request_id": "abc-123",
"variant": "champion",
"conversion": 1,
"revenue": 50.0,
"latency": 0.012
}'curl http://localhost:8000/test-resultResponse:
{
"decision": "continue",
"z_statistic": 1.45,
"boundary": 2.79,
"p_value": 0.147,
"champion_mean": 0.102,
"challenger_mean": 0.115,
"relative_improvement": 0.127,
"recommendation": "Continue collecting data."
}# Deploy 90/10 canary
make deploy-90-10
# Advance to 50/50 validation
make deploy-50-50
# Full rollout (challenger becomes champion)
make deploy-100-0
# Emergency rollback
make rollbackk8s/
├── base/
│ ├── deployment.yaml
│ ├── service.yaml
│ └── kustomization.yaml
└── overlays/
├── 90-10/ # Canary stage
├── 50-50/ # Validation stage
└── 100-0/ # Full rollout
safe-model-deployment/
├── src/
│ ├── router/
│ │ └── traffic_splitter.py # Weighted routing
│ ├── testing/
│ │ ├── sequential_test.py # O'Brien-Fleming
│ │ └── metrics_collector.py # A/B metrics
│ ├── rollback/
│ │ └── auto_rollback.py # Degradation triggers
│ └── api/
│ └── main.py # FastAPI endpoints
├── k8s/
│ ├── base/
│ └── overlays/
├── tests/
├── docker-compose.yml
└── Makefile
- "I implemented O'Brien-Fleming alpha-spending for statistically valid early stopping"
- "Traffic splitting uses consistent hashing for sticky sessions"
- "The system automatically rolls back on >5% metric degradation"
- "Kustomize overlays enable declarative deployment stage management"
- "Guardrail metrics prevent latency/error rate regressions"
- Python 3.11+ / uv
- FastAPI + Pydantic
- NumPy + SciPy (statistics)
- Prometheus metrics
- Kubernetes + Kustomize
MIT