arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01207cs.CVcs.AI

是解码格式,而非扰动:审计视觉语言模型测试时扩展的基于一致性的选择

It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling

Puzhuo Zheng, Hasan Kurban

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现视觉语言模型测试时的选择效果由解码格式而非扰动决定,提出的扰动基选择(Pgs)在控制格式后无显著收益,说明其无法作为有效选择信号。

中文摘要 AI 辅助

测试时扩展通过采样大量候选解并从中选择来提升大语言模型的推理能力,但该方法应用于视觉语言模型(VLMs)时效果不佳:近期研究显示,简单的多数投票优于基于模型自我验证的选择方法,显然是因为在选择层,基于图像的答案与语言先验给出的自信猜测表现一致。一个自然的解决方案是使选择信号成为无法脱离图像计算的信号。我们研究扰动基选择(Pgs),这是一种无标签、无训练的规则,通过模型在输入的标签保留扰动(裁剪、背景遮蔽、轻微光度或几何抖动)下是否重新推导某候选来对其打分;当扰动集为空时,Pgs等价于多数投票。关键问题并非Pgs是否仅击败仅思维链(CoT)的多数投票,而是在控制解码格式和预算后,扰动项是否有额外作用。因此我们引入格式匹配对照组(MatchedCtrl):在原始图像上进行相同的短无CoT采样。在TextVQA、MATH-Vision、MMMU和ViLP四个数据集上,采用Qwen(三次种子的均值)和LLaVA-OneVision的匹配预算选择表,Pgs在TextVQA(Qwen)上较普通多数投票最高提升31.8个百分点,但MatchedCtrl在所有基准(包括需视觉信息的ViLP)中均与Pgs相当或超出(在误差范围内);Qwen的所有类别均未显示出较该对照的显著提升。稳定性差距真实存在且依赖图像(最高达0.48),但无法预测单样本的胜负。该结果是负面且具诊断性的:扰动一致性最多是视觉依赖性的部分诊断,在控制格式后本身并非可用的选择信号;针对仅CoT多数投票报告的收益夸大了此类方法的效果。

英文摘要

Test-time scaling lifts large language model reasoning by sampling many candidate solutions and selecting among them, yet the same recipe transfers poorly to vision-language models (VLMs): recent work shows that simple majority voting beats selection methods built on the model's own self-verification, apparently because at the selection layer an image-grounded answer and a confident guess from the language prior look the same. A natural fix is to make the selection signal one that cannot be computed without the image. We study Perturbation Grounded Selection (Pgs), a label-free, training-free rule that scores each candidate by whether the model re-derives it under label-preserving perturbations of the input (cropping, background masking, mild photometric or geometric jitter); Pgs recovers majority voting when the perturbation set is empty. The decisive question is not whether Pgs beats chain-of-thought only majority voting, but whether the perturbation term adds anything once decoding format and budget are controlled. We therefore introduce a format-matched control (MatchedCtrl): the same short, no-CoT draws spent on the original image. Across TextVQA, MATH-Vision, MMMU, and ViLP, with a Qwen headline (three-seed means) and LLaVA-OneVision coverage in matched-budget selector tables, Pgs appears to beat plain majority voting by up to +31.8 points on TextVQA (Qwen), but MatchedCtrl tracks or exceeds Pgs within noise on every benchmark, including the vision-required ViLP; no Qwen category shows a significant gain over this control. The stability gap is real and image-dependent (up to +0.48), yet does not predict per-instance wins. The result is negative and diagnostic: perturbation consistency is at best a partial diagnostic of visual dependence and, on its own, not a usable selection signal once format is controlled; gains reported against CoT-only majority voting overstate such methods.

↑