推理指令可破坏视觉-语言模型中的答案解码
Reasoning Instructions Can Break Answer Decoding in Vision--Language Models
浏览论文内容
中文总结 AI 辅助
本研究揭示思维链前缀评分会严重破坏视觉-语言模型的多项选择评估,通过实验证明答案信息仍存于隐藏状态但读取失败,并建议避免此类评分方式。
中文摘要 AI 辅助
当评分器附加推理提示但在模型生成任何推理内容之前读取答案标签的logits时,思维链(CoT)指令会扭曲多项选择视觉-语言模型(VLM)评估。我们将此称为CoT前缀评分。在ScienceQA上,Qwen2.5-VL-7B的准确率从80.76%降至45.48%,且在五种选项内容排列中,93.54%的CoT前缀预测选择第一个槽位。条件匹配的线性探针从相同的隐藏状态恢复出78.94%的准确率,而自由生成则恢复75.24%,这表明答案通常在前缀之后仍然存在,而即时读取失败。词汇和层诊断解释了这一不匹配:概率质量向续写标记移动,而答案信息在后期层中仍保持线性可访问。该效应在不同数据集和模型中以不同严重程度重复出现,尽管并非普遍存在。这些结果表明,CoT前缀评分可能将模型知识与评估接口不匹配混淆,应予以避免,除非请求和评分的输出事件是对齐的。
英文摘要
Chain-of-thought (CoT) instructions can distort multiple-choice VLM evaluation when a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. We call this CoT-prefix scoring. On ScienceQA, Qwen2.5-VL-7B drops from 80.76% to 45.48%, and across five option-content permutations 93.54% of CoT-prefix predictions select the first slot. Condition-matched linear probes recover 78.94% from the same hidden states, while free generation restores 75.24%, showing that the answer often survives the prefix and the immediate readout fails. Vocabulary and layer diagnostics explain the mismatch: probability mass moves toward continuation tokens, while answer information remains linearly accessible in late layers. The effect recurs with varying severity across datasets and models, though not universally. These results show that CoT-prefix scoring can confound model knowledge with an evaluation-interface mismatch and should be avoided unless the requested and scored output events are aligned.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。