OCR tier: scanned PDFs + images, local and bomb-proof - #13
Merged
Merged
Conversation
ScannedPdf branch now OCRs embedded page images (lopdf walk, 50-page cap, method pdf-ocr with pages_ocred); new PNG/JPEG claims -> image-ocr; typed failures; models (~12MB) fetch-once to ~/.cache/ocrs (VERITY_OCR_MODEL_DIR override); default-ON cargo feature with clean opt-out; extraction moved to spawn_blocking; png/jpeg mimes in the connector gate + MCP image lane; HONESTY.md: local, printed-text-grade. Real-engine e2e (VERITY_OCR_E2E=1) recognizes rendered text from PNG and scanned PDF. 33 extract tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
R1 launch-blocker fixed: Flate image streams inflate through a hard cap (W*H*channels + predictor slack; flate2 Read::take) — never lopdf's uncapped decompressed_content; a 100KB->100MB bomb now skips typed with bounded memory, probe-test included. R2: DCT dims verified from the codec header BEFORE pixel decode in both lanes. R3: model download gets timeouts, process-wide single-flight, unique temp names, and corrupt-cache self-heal (delete + refetch once). S1: SHA-256 pins for both model files, verified on download and cache load. C1: receipts carry pages_total + pages_skipped_unsupported, and an all-unsupported scan (CCITT/JBIG2/JPX) fails with its own honest reason instead of implying OCR looked. C2: flate-wrapped DCT implemented under the cap. Image lane catch_unwind fenced (typed, no JoinError 500). PNG predictors ported spec-correct (incl. lopdf's Avg precedence bug fixed). 42 extract + 15 engine tests green in BOTH feature configs; real-engine e2e passes; fmt/clippy clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
# Conflicts: # HONESTY.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pure-Rust ocrs (no system deps, no cloud; ~12MB models fetched once, SHA-256-pinned). Scanned PDFs OCR via embedded-image extraction; PNG/JPEG lane; honest receipts (pages_total/pages_skipped, distinct all-unsupported reason). Red-team caught a probe-confirmed decompression-bomb OOM primitive — fixed with capped inflation + pre-decode dimension checks + download timeouts/single-flight/self-heal. 42 extract + 15 engine tests green in both feature configs; real-engine e2e passes.
🤖 Generated with Claude Code