发表机构
ImageCLEF Lab; University of Amsterdam(图像CLEF实验室; 阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究小型视觉语言模型在多语言视觉多项选择题中的测试时缩放,比较多种方法,发现运行条件重要,可解析性是关键,增加解码预算有帮助,复杂方法贡献小,最佳配置在测试集上表现出色,位居排行榜首位。
AI 中文摘要
测试时缩放(TTS)能可靠地提升大语言模型的推理能力,但它能否应用于小型开放视觉语言模型尚不清楚。我们在多语言视觉多项选择题基准EXAMS-V上对此进行研究,比较了自一致性、描述后推理与基于概率推理模型(PRM)引导的束搜索,以及Qwen2.5-VL-7B-Instruct和Qwen3.5-4B上的两种事后选择器。关键在于TTS运行的条件,而非搜索或验证机制。最大因素是可解析性,标准答案提示和引导修复步骤能大幅消除早期提示格式导致的问题。增加解码预算可消除其余问题,提高单链令牌限制从1k到2k可提升3.7个百分点,而增加采样链数量效果不明显。一旦链有足够空间完成,复杂方法贡献不大。最大的提升来自策略模型本身。我们的最佳配置在ImageCLEF 2026测试集上达到84.1%,在视觉多项选择题排行榜上排名第一。
英文摘要
Test-time scaling (TTS) reliably improves reasoning in large language models, but whether it transfers to small open vision-language models remains unclear. We examine this on EXAMS-V, a multilingual visual multiple-choice benchmark, comparing self-consistency, describe-then-reason with PRM-guided beam search, and two post-hoc selectors across Qwen2.5-VL-7B-Instruct and Qwen3.5-4B. What matters is the conditions under which TTS runs, not the search or verification machinery. The largest factor is parseability: an early prompt format left many chains reasoning correctly yet never committing to an answer letter, which a standard answer cue and a guided repair step largely remove. A larger decoding budget removes the rest: raising the per-chain token limit from 1k to 2k recovers 3.7 pp, whereas sampling more chains (8 to 16) adds only 0.15 pp. Once chains have room to finish, elaborate methods contribute little: PRM-guided beam search trails plain self-consistency by 0.39 pp at over eight times the cost, and neither a training-free generative critic nor a trained multimodal PRM beats majority vote across both policies. The largest gain comes instead from the policy model itself (+11.4 pp). Our best configuration reaches 84.1% on the held-out ImageCLEF 2026 test split, ranking first on the Visual MCQ leaderboard.
Comments14 pages, 2 figures, accepted at ImageCLEF 2026