Multi-agent document intelligence platform powered by Claude, LangGraph, and FastAPI
DocuMind AI is a document intelligence system that lets users upload documents and query them through a multi-agent pipeline. A supervisor agent classifies each query and routes it to retrieval, summarization, or direct response, and a corrective critique loop checks generated answers for faithfulness to the source material, retrying with a reformulated query when a check fails.
Every query is fully traced: which agent nodes ran, latency per node, sources cited, tokens used, and cost, all queryable live through Prometheus and Grafana and stored per-request in PostgreSQL.
- Multi-agent pipeline: a supervisor routes each query to retrieval, summarization, or direct response, with a critique node that triggers bounded corrective retries on unfaithful answers
- Hybrid retrieval: dense embeddings, BM25 sparse retrieval, reciprocal rank fusion, and cross-encoder reranking, benchmarked against a dense-only baseline
- JWT authentication: registration, login, token refresh, and logout with Redis-backed token blacklisting
- Full request tracing: every query stores which nodes ran, latency, sources, tokens, and cost, both in PostgreSQL and live in Prometheus/Grafana
- Async ingestion: document upload returns immediately, indexing runs in the background without blocking the API
- Rate limiting: Redis-based per-user request limiting
- Containerized deployment: a single
docker compose upstarts the full stack (API, frontend, PostgreSQL, Redis, Qdrant, Prometheus, Grafana) - Persistent chat UI: a Streamlit interface with saved history, source citations, and an agent trace viewer
- Hardened and tested: timeouts and retries on every external call, schema-validated LLM output, citation validation, a scoped adversarial test set, and a CI regression gate on retrieval and faithfulness quality
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Streamlit UI β
β localhost:8501 β
βββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββββββ
β HTTP
βββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββββββ
β FastAPI Backend β
β Auth Β· Documents Β· Query β
β localhost:8000 β
ββββββββ¬βββββββββββββββ¬βββββββββββββββ¬βββββββββββββββββββββββββ
β β β
ββββββββΌβββββββ βββββββΌβββββββ βββββ βΌβββββββββββββββββββββββββ
β PostgreSQL β β Redis β β LangGraph Agent β
β Users β β Sessions β β β
β Documents β β Rate β β Supervisor β routes query β
β Sessions β β Limits β β Retriever β hybrid+rerank β
β Queries β β Blacklist β β Synthesizerβ generates β
βββββββββββββββ ββββββββββββββ β Critique β faithfulness β
β β + corrective β
β β retry loop β
β ββββββββββββββββ¬ββββββββββββββββ
β β
ββββββββΌββββββββββββββββ βββββββββββββββββΌβββββββββββββββ
β Prometheus/Grafana β β Qdrant β
β Live latency, cost, β β Vector Database β
β retry & error rates β β 1024-dim bge-m3 embeddings β
β localhost:9090/3001 β β (local, per-user filtered) β
ββββββββββββββββββββββββ ββββββββββββββββββββββββββββββββ
User Query
β
βΌ
Input Middleware (pattern-based injection check, blocks obvious attacks)
β
βΌ
Supervisor Node (classifies intent)
β
βββ specific question βββΊ Retriever
βββ summarize βββΊ Summarizer
βββ general βββΊ Synthesizer
β
βΌ
Synthesizer Node
(citations, validated)
β
βΌ
Critique Node
faithful? ββ No βββΊ reformulate query, retry (bounded)
β
Yes
β
βΌ
Final Answer + Sources + Trace + Cost
| Layer | Technology |
|---|---|
| LLM | Claude (claude-haiku-4-5, Anthropic API) |
| Embeddings | BAAI/bge-m3, local and self-hosted (1024 dim) |
| Agent Framework | LangGraph 0.2 |
| API | FastAPI 0.115 + uvicorn |
| Database | PostgreSQL 15 (SQLAlchemy async) |
| Vector DB | Qdrant (dense + BM25 sparse + RRF fusion + cross-encoder rerank) |
| Cache | Redis 7 |
| Frontend | Streamlit 1.40 |
| Observability | Prometheus + Grafana |
| Containerization | Docker + Docker Compose |
| Auth | JWT (python-jose) + bcrypt |
| CI | GitHub Actions, unit tests plus a retrieval/faithfulness regression gate |
- Docker Desktop
- An Anthropic API key (console.anthropic.com); embeddings run locally with no additional key required
git clone https://github.com/bazzal99/Documind-AI.git
cd Documind-AIcp .env.example .env
# Edit .env and add your ANTHROPIC_API_KEYcd infra
docker compose up -d- UI: http://localhost:8501
- API docs: http://localhost:8000/docs
- Health: http://localhost:8000/health
- Grafana: http://localhost:3001 (anonymous viewer access)
- Prometheus: http://localhost:9090
Documind-AI/
βββ backend/
β βββ app/
β β βββ agents/ # LangGraph nodes
β β β βββ graph.py # Agent state, graph definition, routing
β β β βββ supervisor.py # Intent classification and routing
β β β βββ retriever.py # Hybrid (dense + BM25 + RRF) retrieval with reranking
β β β βββ synthesizer.py # Answer generation and citation validation
β β β βββ critique.py # Faithfulness check and corrective retry loop
β β β βββ summarizer.py # Map-reduce summarization
β β βββ api/routes/ # FastAPI endpoints
β β β βββ auth.py # Registration, login, logout
β β β βββ documents.py # Upload, list, delete
β β β βββ query.py # Chat endpoint and input middleware
β β βββ core/ # Config, auth, LLM client, resilience, metrics
β β βββ db/ # Models, session
β β βββ services/ # Document, vector, cache
β βββ tests/ # Unit tests (CI-safe, no live dependencies)
βββ frontend/
β βββ app.py # Streamlit UI
β βββ .streamlit/config.toml # Theme
βββ eval/ # Evaluation scripts behind every number in RESULTS.md
βββ infra/
β βββ docker-compose.yml # Full service stack
β βββ grafana/dashboards/ # Auto-provisioned dashboard
β βββ prometheus/ # Auto-provisioned scrape config
βββ .github/workflows/eval.yml # CI: unit tests and regression gate
βββ docs/images/ # Dashboard screenshot
βββ .env.example
βββ requirements.txt
βββ RESULTS.md # Full evaluation methodology and numbers
βββ README.md
| Method | Endpoint | Description |
|---|---|---|
| POST | /api/v1/auth/register |
Create account |
| POST | /api/v1/auth/login |
Login, returns JWT tokens |
| POST | /api/v1/auth/refresh |
Refresh access token |
| POST | /api/v1/auth/logout |
Blacklist token |
| POST | /api/v1/documents/upload |
Upload and index a document |
| GET | /api/v1/documents/ |
List user documents |
| DELETE | /api/v1/documents/{id} |
Delete document |
| POST | /api/v1/query/ |
Ask a question |
| GET | /api/v1/query/sessions |
List chat sessions |
| GET | /api/v1/query/sessions/{id}/history |
Chat history |
| GET | /health |
Service health check |
| GET | /metrics |
Prometheus metrics |
Full methodology, tables, and caveats for every number below are in RESULTS.md.
Retrieval quality: hybrid retrieval with cross-encoder reranking against a dense-only baseline, over 137 benchmark questions.
| Mode | Recall@5 | MRR | nDCG@10 |
|---|---|---|---|
| Dense only | 0.898 | 0.791 | 0.826 |
| Hybrid + rerank | 0.920 | 0.853 | 0.876 |
Generation quality: faithfulness, relevance, and context precision/recall via an automated judge, over 20 questions per retrieval mode.
| Mode | Faithfulness | Answer Relevance | Context Recall |
|---|---|---|---|
| Dense only | 0.958 | 0.882 | 0.895 |
| Hybrid + rerank | 0.976 | 0.842 | 0.926 |
Hallucination and abstention: measured through the full production pipeline rather than a mocked check.
| Metric | Result |
|---|---|
| Correct Abstention Rate | 0.929 |
| False Abstention Rate | 0.100 |
Corrective retry loop: the critique node's bounded retry mechanism, measured on a 19-question sample.
| Metric | Result |
|---|---|
| Retry Trigger Rate | 0.316 |
| Faithfulness before retry | 0.943 |
| Faithfulness after retry | 0.958 |
Reliability: measured on a 20-question sample.
| Metric | Result |
|---|---|
| Malformed Output Rate | 0.050 |
| Citation Validity Rate | 1.000 |
| External Call Failure Rate | 0 (of ~60 calls) |
Security: a scoped 12-case adversarial test set covering prompt injection, jailbreak attempts, and data extraction.
| Metric | Result |
|---|---|
| Attack Success Rate | 0/12 |
| Cross-user Data Leakage | 0 chunks |
Every request is traced end-to-end and exported to Prometheus and Grafana, backed by the same instrumentation that produces the numbers above.
/metrics (via prometheus-fastapi-instrumentator) exposes default HTTP metrics plus custom series defined in backend/app/core/metrics.py: per-node latency histograms, per-node error counts, retrieval mode distribution, corrective retry and exhaustion counts, token counts, and estimated cost. A trace_id is generated once per request, bound into every node's log line, and stored as an indexed column on the queries table, so a request's full path through the pipeline can be looked up directly. Prometheus and Grafana are provisioned automatically on docker compose up, reachable at localhost:9090 and localhost:3001.
Panels, left to right, top to bottom: node latency p50/p95/p99, node error rate, corrective retry rate, retrieval mode distribution, tokens consumed, cost per request, spend in the last 24 hours, and cumulative spend. Full latency-tracking methodology is in RESULTS.md.
The pipeline hardens against the failure modes that occur in production. Every external dependency has explicit timeouts and bounded retries: Claude through the Anthropic SDK's native exponential backoff, Qdrant and Redis through a jittered retry helper, since neither client retries on its own. Every attempt is tracked in Prometheus per service. The critique node's faithfulness verdict is parsed through a Pydantic schema rather than a bare dictionary lookup, so a malformed response is caught and counted rather than silently defaulting to "faithful." Every citation in a generated answer is validated against the sources actually retrieved for that query.
Full numbers in RESULTS.md.
A scoped adversarial test set, not a comprehensive red-team suite, covering prompt injection, jailbreak attempts, and data-extraction attempts, plus a direct cross-user isolation test. Pattern-based input middleware blocks obviously-phrased attacks before they reach the agent pipeline. Every retrieval is scoped by a per-user filter, confirmed with an explicit isolation test rather than assumed. Emails and phone numbers are redacted from application logs.
Full numbers in RESULTS.md.
| Area | Script |
|---|---|
| Retrieval quality | eval.build_dataset, eval.seed_corpus, eval.run_retrieval_ablation |
| Generation quality | eval.run_generation_eval |
| Hallucination and abstention | eval.week3_hallucination, eval.check_abstention |
| Corrective retry loop | eval.week4_corrective_loop |
| Reliability | eval.week6_reliability |
| Security | eval.security_tests |
| CI regression gate | eval.ci_eval |
.github/workflows/eval.yml runs on every push and pull request to main. Unit tests always run, with no external dependencies. A regression-eval job runs a fixed 15-question subset through retrieval and generation and fails the build if Recall@5 or faithfulness drops more than 0.05 from the recorded baseline (eval/baseline_scores.json); it is opt-in via an ANTHROPIC_API_KEY repository secret plus an ANTHROPIC_API_KEY_CONFIGURED=true repository variable, since it spends API credits on every run.
{
"answer": "The methodology uses a ConvLSTM2D architecture trained on 2,000 videos...",
"sources": [
{
"filename": "research_paper.pdf",
"relevance_score": 0.94
}
],
"nodes_invoked": ["supervisor", "retriever", "synthesizer", "critique"],
"agent_trace": [
{"node": "supervisor", "route": "retriever", "latency_ms": 702},
{"node": "retriever", "chunks_found": 5, "latency_ms": 366},
{"node": "synthesizer", "invalid_citations": 0, "latency_ms": 2100},
{"node": "critique", "faithful": true, "latency_ms": 615}
],
"latency_ms": 3800,
"tokens_used": 1240,
"cost_usd": 0.00187
}Claude is the only paid dependency; embeddings run locally and every other service is self-hosted. Per-request cost is tracked live and visible on the Grafana dashboard rather than estimated after the fact. At current pricing, a typical query costs well under a cent.
Mohammad Bazzal, ML Engineer, PhD in Telecommunications, 3x IEEE Author
MIT License.
