DocuChat is a Retrieval-Augmented Generation (RAG) web app: upload a document (.txt or .pdf) and ask questions about it in a ChatGPT-style chat interface. It chunks and embeds the document with HuggingFace embeddings, stores them in an in-memory Chroma vector store, and generates grounded, streamed answers using Azure OpenAI — served through a Flask backend.
Retrieval quality isn't just claimed — it's measured. eval/eval_retrieval.py is a 40-question benchmark (built from the actual sample corpus) that compares naive vector search against hybrid BM25+vector search, both at the retrieval level and end-to-end through real generated answers. See Retrieval evaluation below.
Browser <-- stream --> Flask (app/routes.py)
|
+-- app/rag.py load + split uploaded doc
+-- Chroma (in-memory) per-session vector store
+-- app/session_store.py per-session chat history (in-memory)
+-- Azure OpenAI answer generation (streamed)
Each browser session (a signed cookie, no login) gets its own isolated Chroma collection and chat history — uploads and conversations never leak across users. See Production notes for the scaling trade-off this implies.
.
├── app/ # Flask app (the product)
│ ├── __init__.py # create_app() factory
│ ├── config.py # env-driven configuration
│ ├── rag.py # document loading/splitting/vectorstore helpers
│ ├── session_store.py # per-session state (in-memory)
│ ├── routes.py # / , /api/status, /api/upload, /api/chat, /api/reset
│ ├── static/, templates/ # chat UI
│ └── uploads/ # scratch space for uploads (deleted after processing)
├── run.py # local dev entrypoint
├── wsgi.py # production entrypoint (gunicorn/waitress)
├── eval/eval_retrieval.py # retrieval + end-to-end accuracy benchmark
├── data/sample_docs/ # sample corpus (5 company profiles) used by eval/
├── tests/ # pytest suite (no Azure calls required)
├── Dockerfile
└── .github/workflows/ci.yml # lint + test on push/PR
- Python 3.10+
- An Azure OpenAI resource with a chat model deployed (e.g.
gpt-4o-mini,gpt-4o) — see Setting up Azure OpenAI
-
Clone and enter the repo, create a virtual environment
git clone <your-repo-url> cd DocuChat python -m venv venv # Windows venv\Scripts\activate # macOS/Linux source venv/bin/activate
-
Install dependencies
pip install -r requirements.txt
-
Configure environment variables
cp .env.example .env
Then edit
.env:Variable Description AZURE_OPENAI_API_KEYFrom Azure Portal → your Azure OpenAI resource → Keys and Endpoint AZURE_OPENAI_ENDPOINTThe bare resource endpoint, e.g. https://<your-resource-name>.openai.azure.com/(not the Azure AI Foundry project endpoint — they're different and not interchangeable)AZURE_OPENAI_API_VERSIONDefaults to 2024-05-01-preview; only change if your deployment needs a different versionAZURE_OPENAI_DEPLOYMENT_NAMEThe deployment name you gave your chat model in Azure AI Foundry → Deployments (can differ from the underlying model name) FLASK_SECRET_KEYAny random string, used to sign session cookies. Generate one with python -c "import secrets; print(secrets.token_hex(32))"
- Create an Azure OpenAI resource in the Azure Portal.
- In Azure AI Foundry, open your resource's project → Deployments → deploy a chat-completion-capable model (e.g.
gpt-4o-mini).- Reasoning/Codex-family models often only support the Responses API, not Chat Completions — stick to a standard GPT chat model.
- Grab the API key and endpoint from Azure Portal → your resource → Keys and Endpoint.
Local (dev):
python run.pyOpen http://127.0.0.1:5000, upload a .txt or .pdf file (try one from data/sample_docs/), and start asking questions.
Local (production-like WSGI server):
# Linux/macOS
gunicorn -w 1 --threads 4 -b 0.0.0.0:5000 wsgi:app
# Windows
waitress-serve --host=0.0.0.0 --port=5000 wsgi:appDocker:
docker build -t docuchat .
docker run --env-file .env -p 5000:5000 docuchateval/eval_retrieval.py grades retrieval and generation against a 40-question, fact-grounded test set built from data/sample_docs/ (e.g. "Who founded Nvidia and when?" → the retrieved/generated text must contain "Jensen Huang" and "1993"). It builds data/chroma_db from data/sample_docs/ automatically on first run if it doesn't exist yet.
python eval/eval_retrieval.py # retrieval-only, free (no LLM calls)
python eval/eval_retrieval.py --end-to-end # also grades real generated answers (uses Azure OpenAI)Measured results on this corpus:
| Metric | Naive vector search | Hybrid (BM25 + vector) |
|---|---|---|
| Retrieval Hit Rate@5 | 67.5% | 80.0% |
| Retrieval MRR@5 | 0.537 | 0.672 |
| End-to-end answer accuracy | 70.0% | 80.0% |
The webapp itself currently uses naive vector search (db.as_retriever(search_kwargs={"k": 5}) in app/routes.py) — the hybrid strategy is validated here and is the natural next upgrade for app/rag.py.
pip install -r requirements-dev.txt
pytest
ruff check .tests/test_routes.py exercises the full upload → chat → reset flow through Flask's test client with a fake LLM injected via create_app(llm=..., embedding_model=...) — no Azure credentials or network calls needed, so it runs the same locally and in CI (.github/workflows/ci.yml).
This is a portfolio-scale deployment, and it's honest about where that shows:
- Session state is in-memory (
app/session_store.py): uploaded documents and chat history live in a process-local dict keyed by session cookie. This is why the Docker image and the gunicorn command both run 1 worker — splitting requests across multiple processes would make sessions "disappear" depending on which worker handled a given request. Threads still provide concurrency within that worker. Scaling beyond one worker would mean moving session + vector-store state to a shared backend (e.g. Redis + a hosted vector DB) instead. - Uploaded files are session-scoped and ephemeral: saved to
app/uploads/only long enough to chunk and embed, then deleted; nothing is persisted to disk after that. - Retrieval is currently naive vector search in the live app, even though the hybrid strategy measurably outperforms it (see above) — swapping
app/rag.py's retriever is the natural next step, not yet wired into the webapp.
.envis git-ignored — never commit real API keys. Use.env.exampleas the template.data/chroma_db/(the eval harness's persisted vector store) andapp/uploads/are git-ignored.