Hybrid retrieval (BM25 + MiniLM dense embeddings) with cross-encoder reranking over Docling-chunked financial reports (10-K/10-Q, annual reports, earnings releases).
PDF -> Docling (layout, tables, headings) -> metadata extraction
-> HybridChunker chunking -> [BM25 index] + [MiniLM/ChromaDB index]
-> RRF fusion -> cross-encoder rerank
-> context construction -> LLM (Groq or HF Inference API) -> cited answer
-> FastAPI -> Streamlit UI
- Docling for parsing + chunking: Docling produces structurally-aware
parsing (real reading order,
TableFormertable structure, real headings) instead of raw character extraction. ItsHybridChunkeralready respects those structural boundaries (won't split a table or cut mid-sentence) and carries page/heading provenance per chunk, so chunks are indexed and sent to the LLM as a single tier — no parent/child splitting needed to compensate for a naive splitter. - Hybrid BM25 + dense: BM25 catches exact terms (ticker symbols, line-item names, dollar figures) that embeddings can blur; dense embeddings catch paraphrases ("how did profitability trend" -> "operating margin"). Fused with Reciprocal Rank Fusion (rank-based, not raw score averaging — the two scales aren't comparable). Dense side uses ChromaDB as a persistent local vector store (MiniLM embeddings computed via SentenceTransformers, added to Chroma as raw vectors).
- Cross-encoder rerank: bi-encoder cosine similarity is fast but coarse; a cross-encoder jointly attends over query+doc for much higher precision on the ~15 shortlisted candidates.
- Citation via numbered context blocks: the LLM is instructed to
cite
[n]markers that map deterministically back to source chunks — citations aren't generated freeform by the model, so they can't drift from the real page/section. Citation text leads with section- page (from Docling, reliable) rather than company/year/quarter (from filename/cover-page heuristics, which stay unreliable).
- Pluggable LLM provider:
app/llm_client.pyis a small factory producing a uniformChatClientover either Groq or the HF Inference API. The generator and the RAGAS judge each pick their own provider/model independently viasettings.generator_provider/generator_model_idandjudge_provider/judge_model_id— e.g. run generation on Groq and judge with a different model, with no code changes.
python -m venv venv && source venv/bin/activate
pip install -r requirements.txt
cp .env.example .env # add GROQ_API_KEY and/or HF_API_TOKEN# 1. Drop PDFs into data/raw/, then build the index
python -m app.ingest
# 2. Run the API
uvicorn app.main:app --reload
# 3. Query
curl -X POST localhost:8000/query \
-H "Content-Type: application/json" \
-d '{"question": "What was total revenue in Q3?"}'
# 4. (optional) Streamlit UI
streamlit run ui/streamlit_app.pydocker compose run --rm ingest # one-off: build the index
docker compose up # backend (8000) + UI (8501) + redis# Fill in evaluation/eval_dataset.json with real Q/A pairs first
python -m evaluation.eval_ragasReports faithfulness, answer relevancy, context precision, and context recall — faithfulness is the one to watch most closely for financial data, since it directly measures hallucinated figures.