Skip to content

Repository files navigation

Financial Reports Hybrid RAG

Hybrid retrieval (BM25 + MiniLM dense embeddings) with cross-encoder reranking over Docling-chunked financial reports (10-K/10-Q, annual reports, earnings releases).

Architecture

PDF -> Docling (layout, tables, headings) -> metadata extraction
     -> HybridChunker chunking -> [BM25 index] + [MiniLM/ChromaDB index]
     -> RRF fusion -> cross-encoder rerank
     -> context construction -> LLM (Groq or HF Inference API) -> cited answer
     -> FastAPI -> Streamlit UI

Why these design choices

  • Docling for parsing + chunking: Docling produces structurally-aware parsing (real reading order, TableFormer table structure, real headings) instead of raw character extraction. Its HybridChunker already respects those structural boundaries (won't split a table or cut mid-sentence) and carries page/heading provenance per chunk, so chunks are indexed and sent to the LLM as a single tier — no parent/child splitting needed to compensate for a naive splitter.
  • Hybrid BM25 + dense: BM25 catches exact terms (ticker symbols, line-item names, dollar figures) that embeddings can blur; dense embeddings catch paraphrases ("how did profitability trend" -> "operating margin"). Fused with Reciprocal Rank Fusion (rank-based, not raw score averaging — the two scales aren't comparable). Dense side uses ChromaDB as a persistent local vector store (MiniLM embeddings computed via SentenceTransformers, added to Chroma as raw vectors).
  • Cross-encoder rerank: bi-encoder cosine similarity is fast but coarse; a cross-encoder jointly attends over query+doc for much higher precision on the ~15 shortlisted candidates.
  • Citation via numbered context blocks: the LLM is instructed to cite [n] markers that map deterministically back to source chunks — citations aren't generated freeform by the model, so they can't drift from the real page/section. Citation text leads with section
    • page (from Docling, reliable) rather than company/year/quarter (from filename/cover-page heuristics, which stay unreliable).
  • Pluggable LLM provider: app/llm_client.py is a small factory producing a uniform ChatClient over either Groq or the HF Inference API. The generator and the RAGAS judge each pick their own provider/model independently via settings.generator_provider/ generator_model_id and judge_provider/judge_model_id — e.g. run generation on Groq and judge with a different model, with no code changes.

Setup

python -m venv venv && source venv/bin/activate
pip install -r requirements.txt
cp .env.example .env   # add GROQ_API_KEY and/or HF_API_TOKEN

Usage

# 1. Drop PDFs into data/raw/, then build the index
python -m app.ingest

# 2. Run the API
uvicorn app.main:app --reload

# 3. Query
curl -X POST localhost:8000/query \
  -H "Content-Type: application/json" \
  -d '{"question": "What was total revenue in Q3?"}'

# 4. (optional) Streamlit UI
streamlit run ui/streamlit_app.py

Docker

docker compose run --rm ingest   # one-off: build the index
docker compose up                # backend (8000) + UI (8501) + redis

Evaluation

# Fill in evaluation/eval_dataset.json with real Q/A pairs first
python -m evaluation.eval_ragas

Reports faithfulness, answer relevancy, context precision, and context recall — faithfulness is the one to watch most closely for financial data, since it directly measures hallucinated figures.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages