Compare a source PDF with a translation (or reprint) and mark layout mismatches: field alignment, lines, colors, fonts, sizes.
Layout extraction uses the Huridocs Docker image when it is up. Without it, the pipeline uses JSON next to each PDF (samples/ includes fixtures).
First image build downloads PyTorch — expect several minutes.
python samples/generate_samples.py # if samples/*.pdf are missing
docker compose up --buildAnnotated PDFs land in Result/ on the host.
Skip Huridocs and use sidecar JSON only:
docker compose run --rm -e LAYOUT_SERVICE_URL= app python -m pdfqaLinux: sudo apt-get install poppler-utils
python -m venv .venv
# Windows: .venv\Scripts\activate
pip install -r requirements.txt
pip install -e .
python samples/generate_samples.py
python -m pdfqaOptional layout service:
docker run --rm --name pdf-layout -p 5060:5060 --entrypoint ./start.sh huridocs/pdf-document-layout-analysis:v0.0.21Override the URL: LAYOUT_SERVICE_URL=http://127.0.0.1:5060
from pdfqa.pipeline import compare_pdfs
from pdfqa.cli import configure_and_run
compare_pdfs("samples/form_en.pdf", "samples/form_es.pdf")
configure_and_run(pdf1="a.pdf", pdf2="b.pdf", color_threshold=80)Thresholds: src/pdfqa/config.py
src/pdfqa/
pipeline.py # orchestrator
cli.py # config overrides
layout/ # JSON sort, merge, semantic field match
visual/ # lines, colors, fonts, crops, annotations
util/
samples/ # synthetic EN/ES forms
Outputs: Result/doc1_annotated.pdf, Result/doc2_annotated.pdf
Do not commit third-party insurance forms. With all gratitute toward my friend Tony, whom I was too young to understand