Skip to content

Repository files navigation

Incident Ledger

CI

Evidence-first incident operations, with a hard boundary around human authority.

Incident Ledger is a multi-agent investigation workspace for synthetic operational incidents. Three specialist agents analyze signals, test the incident hypothesis, and prepare a bounded response plan; a fourth agent synthesizes their reports. Evidence-ledger entries carry verdicts and evidence IDs, while every state-changing proposal remains behind a human approval gate.

The included scenario is entirely synthetic. It contains no company logs, endpoints, credentials, or operational data.

Incident Ledger architecture

Why it exists

During an incident, teams do not need another chatbot producing an eloquent guess. They need a compact, auditable answer to three questions:

  1. What do we actually know?
  2. Which conclusion is inference rather than evidence?
  3. What can an agent prepare without being allowed to change a live system?

Incident Ledger makes those distinctions part of the product, not a footnote in a prompt.

What is implemented

  • An asymmetric incident workspace with a progressive evidence ledger, agent roster, source-level citations, and explicit authority gate.
  • A deterministic synthetic demonstration that remains usable without cloud credentials.
  • A Python service for structured incident analysis, designed for Google Agent Development Kit and Gemini 3.5 Flash.
  • Role-separated signal-analysis, evidence-verification, and runbook-planning agents.
  • Two explicit read-only ADK tools that list and retrieve only synthetic:// telemetry; all other targets are rejected.
  • A safety invariant: read-only checks may be PROPOSED; every state-changing action is forced to REQUIRES_HUMAN_APPROVAL.
  • Honest google_adk / deterministic_demo response labels, including provider-failure fallback handling.
  • Cloud Run packaging and an optional hash-linked, append-only Firestore audit sink.
  • Default cost guards of 10 runs per UTC day and 50 runs for the deployment lifetime, enforced globally with Firestore transactions and fail-closed on errors.
  • Tests covering response shape, safe fallback behavior, and the authority boundary.

Run the workspace

Requires Node.js 22.13 or newer.

npm install
npm run dev

Open http://localhost:3000. With no API URL configured, the interface uses the deterministic browser demonstration.

To connect a deployed service, copy .env.example to .env.local and set:

NEXT_PUBLIC_AGENT_API_URL=https://incident-ledger-agent-414211538718.asia-east1.run.app

The hosted judge demo is connected to this public, cost-capped Cloud Run API. Its health endpoint is available at /health.

Run the agent service

The service is self-contained under agent_service/. It can run without a Gemini key in deterministic demo mode; cloud credentials activate the ADK path.

cd agent_service
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn main:app --reload

See the service guide for environment variables, endpoints, tests, and Cloud Run deployment.

Safety model

Layer Allowed Never automatic
Signal analyst Read and summarize synthetic telemetry Modify alerts or services
Evidence verifier Cross-check sources and label uncertainty Invent or conceal evidence
Runbook planner Prepare a bounded, reversible proposal Restart, deploy, roll back, or delete
Review UI Mark evidence reviewed in the browser Persist approval or authorize or execute a change

This repository is a competition prototype, not a production incident-management system. Production connectors would require tenant isolation, authentication, authorization, secret management, rate limits, and an organization-specific approval policy.

Competition fit

Incident Ledger is designed for the Taskmaster category, with a strong fit for the Individual/Hobbyist and Best Architectural Design prizes in Google's All Things Agentic Hackathon. The new work in this repository combines Google ADK orchestration, Gemini reasoning, a Cloud Run deployment path, source-grounded outputs, and a visible human-in-the-loop boundary.

The same safety-first core can later be adapted to AWS Strands or a consent-based phone agent without using private operational data.

Submission assets

License

MIT

About

Evidence-first multi-agent incident operations with a hard human authority gate.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages