Reproducible framework for document-centric VLM robustness and claim-faithfulness evaluation.
-
Updated
Jul 18, 2026 - Python
Reproducible framework for document-centric VLM robustness and claim-faithfulness evaluation.
A specialized benchmarking framework for evaluating state-of-the-art Medical Vision-Language Models (VLMs) on clinical imaging datasets, optimized for resource-constrained environments and regional disease prevalence.
Do frontier VLMs (Gemini/GPT/Claude) infer a robot’s goal from partial motion? Goal-inference benchmark for robot-motion legibility - MS thesis code.
Description: Domain-shift failure analysis of LLaVA-1.5-7B and InternVL2-8B across photos, charts, medical, and screenshots. 945 probes, 6 failure categories. Chart OCR: LLaVA 0% vs InternVL2 71%. Domain-dependent failure confirmed via chi-square (p<0.0001).
Add a description, image, and links to the vlm-evaluation topic page so that developers can more easily learn about it.
To associate your repository with the vlm-evaluation topic, visit your repo's landing page and select "manage topics."