Independent re-measurement of two ZDTaichu5.0-9B model-card numbers on the full sets: CV-Bench 87.26% [85.93, 88.51] against a card of 86.82, MathVista 82.40% [79.90, 84.71] against 84.50 — and a verdict that turns on how 37 truncated generations are counted.
benchmark reproducibility multimodal evaluation-protocol vision-language-model vllm qwen3 gated-deltanet mathvista cv-bench
-
Updated
Sep 23, 2026 - Python