跨格式共享图片处理管线 - #4
Merged
Merged
Conversation
新增 src/opendocs/vision/images.py 作为共享图片准入、标准化与长图切片模块, 统一替换各解析器原先分散的图片前处理逻辑。核心能力包括: - 透明图片根据前景亮度自适应选择黑/白对比背景,避免语义内容消失 - 高置信度过滤空白、近乎全透明及稀疏装饰图标(结合 alpha coverage、 连通组件、颜色/边缘复杂度、placement area、alt text) - 独立上传小图片不因尺寸硬过滤;无意义图片由模型返回 elements=[] - 长图按约 10% 重叠纵向切片、顺序合并、bbox 映射,超过 32 tile 抛 LimitExceededError;部分 tile 失败保留成功内容 - Office 同 digest 只分析一次,按 (page_number, source_index) 独立准入 - PDF hybrid crop 接入共享准入与切片,tile→crop→page 两级 bbox 映射 - PDF 多 region、Office 多图片恢复并发模型调用 - native preparation / 模型 fatal 错误直接重抛,不被部分成功掩盖 - 加固临时文件清理(suppress OSError)和 native wire workspace 越界保护 - 图片像素预算从 80MP 收紧到 40MP Constraint: 不改变公共 API 和现有错误契约; 不读取/提交私有 corpus Rejected: 不单独改造 PDF renderer(仍受限于最长边 2048px 渲染) Confidence: 三轮独立代码审查均无 Blocker/High; DeepEcho PPTX 真实文档验证通过 Scope-risk: 涉及共享图片模块和 image/office/pdf 三个解析器; 旧路径(无视觉模型)的 skip 语义可能产生少量新增 warning; 不影响已启用视觉模型的解析结果 Tested: | 目标图片管线 84 项回归通过,全量 591 passed / 9 skipped; Ruff check、Ruff format check、ty check src tests、git diff --check、 uv build 全部通过 Not-tested: PDF renderer 超长页自适应高分辨率渲染; 私有 corpus gate (tests/test_acceptance_corpus.py --corpus-dir=@Local)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
目标
为独立图片、DOCX、PPTX、PDF 引入跨格式共享图片处理管线,统一准入、透明图标准化、装饰图过滤和长图切片逻辑。
核心变更
新增
src/opendocs/vision/images.pyelements=[]LimitExceededError修改
src/opendocs/parsers/image.py修改
src/opendocs/parsers/office/parser.py/merge.py(page_number, source_index)独立准入修改
src/opendocs/parsers/pdf/parser.py修改
src/opendocs/vision/prompts.py修改
src/opendocs/vision/litellm.pyelements=[]验证
残余限制