Reliability Decision Control Plane
A prototype demonstrating how AI can safely participate in production operations without removing accountability.
Instead of letting AI execute actions directly, the system separates recommendation, trust, and execution using deterministic guardrails, confidence thresholds, and human-readable explanations.
⸻
Why this exists
During incidents, outages often worsen not because engineers lack knowledge — but because humans must interpret multiple conflicting signals under pressure.
This project explores a safer model:
AI suggests → Policy decides → Automation acts → Humans audit
The goal is reducing operational decision latency without introducing unsafe autonomy.
⸻
What the system does
Given telemetry signals, the system determines one of four actions: • RESTART_SERVICE • WAIT_AND_OBSERVE • ESCALATE_TO_HUMAN • DO_NOT_RESTART
Then applies governance rules before allowing automation.
⸻
Architecture
- Deterministic Safety Layer
Handles known failure patterns immediately: • Dependency failures → never restart • Restart loops → escalate • Memory leaks → restart • Noise signals → observe
- LLM Reasoning Layer
Only ambiguous situations go to the model for contextual evaluation.
- Governance Layer
Automation allowed only if confidence ≥ environment threshold
Environment Required Confidence
prod 0.90 staging 0.70 dev 0.50
- Explainability Layer
Generates human-readable justification for auditability.
⸻
API Endpoints
Decision (machine action)
POST /decide
{ "alert": "Memory usage 99% on worker", "errors": "OutOfMemory exception", "dependency": "none", "previous_restarts": 0, "environment": "staging" }
Response: { "decision": "RESTART_SERVICE", "confidence": 0.85, "automation_allowed": true, "required_confidence": 0.7 }
Explanation (human understanding)
POST /explain
{ "alert": "Payments API latency spike", "errors": "Database timeout errors increasing", "dependency": "Primary DB latency 2500ms", "previous_restarts": 1, "environment": "prod" }
{ "decision": { "decision": "DO_NOT_RESTART", "confidence": 0.95 }, "policy": { "environment": "prod", "required_confidence": 0.9 }, "explanation": { "summary": "Restart avoided because dependency is unhealthy", "risk": "Restart would not recover service and may extend outage" } }
Key Idea
This is not an AI chatbot.
It is a reliability decision model showing how AI can assist operations without creating uncontrolled automation.
Core principles: • AI recommendations are not automatically trusted • Production safety policies override model confidence • Automation maturity can increase gradually • Every decision must remain explainable
⸻
Tech Stack • FastAPI • Local LLM (Ollama) • Python
⸻
Running locally
Start LLM:
ollama serve ollama run llama3
uvicorn app:app --reload --port 8000
Concept Demonstrated
A practical pattern for regulated environments:
Progressively Trusted Automation
Not zero-touch operations. Not human-only operations.
A controlled path between the two. :::