This guide covers repository layout, development commands, testing, and implementation notes.
app/api/ FastAPI routes
app/core/ settings, logging, typed errors
app/storage/ file uploads and SQLite job store
app/parsers/ PDF and PPTX parsers
app/generation/ chunking, prompts, LLM clients (Ollama + Google), factory, orchestration
app/pipeline.py parse -> generate -> persist pipeline
app/worker.py durable worker loop
tests/ unit, API, parser, and LLM contract tests
scripts/ benchmark tooling
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtRun the API:
uvicorn app.main:app --reloadRun tests and lint:
pytest
ruff check .Run a benchmark:
python scripts/benchmark.py path/to/document.pdfBenchmark records are appended to storage/benchmarks.jsonl.
POST /upload
|
v
FileUploadManager
|
v
SQLite JobStore -----> GET /status/{document_id}
|
v
Worker loop
|
+--> PDFParser -> sections + images + diagnostics
+--> PPTXParser -> slides + images + notes + diagnostics
|
v
Extractive evidence prefilter
|
v
LocalLLMOrchestrator -> LLM client (Ollama or Google GenAI via factory)
|
v
FinalDocument JSON -> GET /result/{document_id}
Expected failures use typed errors from app/core/errors.py and are persisted into job state:
unsupported_file_typefile_too_largeempty_fileencrypted_pdfcorrupt_documentollama_unavailableollama_timeoutgoogle_api_errorgoogle_timeoutllm_provider_invalidllm_json_invalidresult_not_ready
Clients inspect failures through /status/{document_id}.
Tests use pytest, pytest-asyncio, and pytest-cov. LLM behavior is tested with fake clients (both Ollama and Google) so CI stays deterministic. The shared response normalization in _response.py has dedicated tests covering JSON parsing, markdown stripping, and schema coercion. Live model behavior belongs in benchmark runs.
Fast mode reduces LLM latency by parsing native text first, building a bounded extractive evidence set, and running three JSON prompts in parallel.
- Ollama: Concurrency is memory-sensitive; increasing
OLLAMA_NUM_PARALLELcan improve throughput but also multiplies context memory use. - Google GenAI: Latency depends on network and API quotas. The
gemini-2.0-flashmodel is recommended for speed and cost-effectiveness.