Skip to content

About

A prototype reliability control plane using AI reasoning with deterministic guardrails and confidence-based execution governance.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

Reliability Decision Control Plane

A prototype demonstrating how AI can safely participate in production operations without removing accountability.

Instead of letting AI execute actions directly, the system separates recommendation, trust, and execution using deterministic guardrails, confidence thresholds, and human-readable explanations.

⸻

Why this exists

During incidents, outages often worsen not because engineers lack knowledge — but because humans must interpret multiple conflicting signals under pressure.

This project explores a safer model:

AI suggests → Policy decides → Automation acts → Humans audit

The goal is reducing operational decision latency without introducing unsafe autonomy.

⸻

What the system does

Given telemetry signals, the system determines one of four actions: • RESTART_SERVICE • WAIT_AND_OBSERVE • ESCALATE_TO_HUMAN • DO_NOT_RESTART

Then applies governance rules before allowing automation.

⸻

Architecture

  1. Deterministic Safety Layer

Handles known failure patterns immediately: • Dependency failures → never restart • Restart loops → escalate • Memory leaks → restart • Noise signals → observe

  1. LLM Reasoning Layer

Only ambiguous situations go to the model for contextual evaluation.

  1. Governance Layer

Automation allowed only if confidence ≥ environment threshold

Environment Required Confidence

prod 0.90 staging 0.70 dev 0.50

  1. Explainability Layer

Generates human-readable justification for auditability.

⸻

API Endpoints

Decision (machine action)

POST /decide

{ "alert": "Memory usage 99% on worker", "errors": "OutOfMemory exception", "dependency": "none", "previous_restarts": 0, "environment": "staging" }

Response: { "decision": "RESTART_SERVICE", "confidence": 0.85, "automation_allowed": true, "required_confidence": 0.7 }

Explanation (human understanding)

POST /explain

{ "alert": "Payments API latency spike", "errors": "Database timeout errors increasing", "dependency": "Primary DB latency 2500ms", "previous_restarts": 1, "environment": "prod" }

{ "decision": { "decision": "DO_NOT_RESTART", "confidence": 0.95 }, "policy": { "environment": "prod", "required_confidence": 0.9 }, "explanation": { "summary": "Restart avoided because dependency is unhealthy", "risk": "Restart would not recover service and may extend outage" } }

Key Idea

This is not an AI chatbot.

It is a reliability decision model showing how AI can assist operations without creating uncontrolled automation.

Core principles: • AI recommendations are not automatically trusted • Production safety policies override model confidence • Automation maturity can increase gradually • Every decision must remain explainable

⸻

Tech Stack • FastAPI • Local LLM (Ollama) • Python

⸻

Running locally

Start LLM:

ollama serve ollama run llama3

uvicorn app:app --reload --port 8000

http://localhost:8000/docs

Concept Demonstrated

A practical pattern for regulated environments:

Progressively Trusted Automation

Not zero-touch operations. Not human-only operations.

A controlled path between the two. :::

About

A prototype reliability control plane using AI reasoning with deterministic guardrails and confidence-based execution governance.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages