Skip to content

OCR tier: scanned PDFs + images, local and bomb-proof - #13

Merged
athenanewsapi merged 3 commits into
mainfrom
ocr-extraction
Aug 5, 2026
Merged

athenanewsapi merged 3 commits into
mainfrom
ocr-extraction

Conversation

@athenanewsapi

Copy link
Copy Markdown
Contributor

Pure-Rust ocrs (no system deps, no cloud; ~12MB models fetched once, SHA-256-pinned). Scanned PDFs OCR via embedded-image extraction; PNG/JPEG lane; honest receipts (pages_total/pages_skipped, distinct all-unsupported reason). Red-team caught a probe-confirmed decompression-bomb OOM primitive — fixed with capped inflation + pre-decode dimension checks + download timeouts/single-flight/self-heal. 42 extract + 15 engine tests green in both feature configs; real-engine e2e passes.

🤖 Generated with Claude Code

athenanewsapi and others added 3 commits August 4, 2026 21:41
ScannedPdf branch now OCRs embedded page images (lopdf walk, 50-page cap,
method pdf-ocr with pages_ocred); new PNG/JPEG claims -> image-ocr; typed
failures; models (~12MB) fetch-once to ~/.cache/ocrs (VERITY_OCR_MODEL_DIR
override); default-ON cargo feature with clean opt-out; extraction moved to
spawn_blocking; png/jpeg mimes in the connector gate + MCP image lane;
HONESTY.md: local, printed-text-grade. Real-engine e2e (VERITY_OCR_E2E=1)
recognizes rendered text from PNG and scanned PDF. 33 extract tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
R1 launch-blocker fixed: Flate image streams inflate through a hard cap
(W*H*channels + predictor slack; flate2 Read::take) — never lopdf's uncapped
decompressed_content; a 100KB->100MB bomb now skips typed with bounded memory,
probe-test included. R2: DCT dims verified from the codec header BEFORE pixel
decode in both lanes. R3: model download gets timeouts, process-wide
single-flight, unique temp names, and corrupt-cache self-heal (delete +
refetch once). S1: SHA-256 pins for both model files, verified on download
and cache load. C1: receipts carry pages_total + pages_skipped_unsupported,
and an all-unsupported scan (CCITT/JBIG2/JPX) fails with its own honest
reason instead of implying OCR looked. C2: flate-wrapped DCT implemented
under the cap. Image lane catch_unwind fenced (typed, no JoinError 500).
PNG predictors ported spec-correct (incl. lopdf's Avg precedence bug fixed).

42 extract + 15 engine tests green in BOTH feature configs; real-engine e2e
passes; fmt/clippy clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
# Conflicts:
#	HONESTY.md
@athenanewsapi
athenanewsapi merged commit 1db08a8 into main Aug 5, 2026
7 of 8 checks passed
@athenanewsapi
athenanewsapi deleted the ocr-extraction branch August 5, 2026 15:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant