TRAPSBench:视觉-语言模型可编码认知约束却无法表达
TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint
浏览论文内容
中文总结 AI 辅助
研究针对视觉-语言模型的认知约束表达问题,构建TRAPSBench基准并提出PECS指标,发现模型可感知不确定性却难表达,瓶颈在表达,需输出阶段干预。
中文摘要 AI 辅助
当视觉证据被遮挡或混乱时,模型应弃权(不执行)。本文中,我们表明视觉-语言模型(VLMs)可在内部区分何时需要弃权,但仍无法将其表达出来。我们推出TRAPSBench,这是一个程序生成的视频基准,包含1404个匹配的物理对,其中单个目标变化会使视觉证据无法确定结果。此外,我们推出惩罚性认知校准分数(PECS),这是一种新的鲁棒指标,要求模型在结果可知时正确回答,且在结果不可知时弃权。在涵盖五个系列的16个VLMs中,自发约束表现不佳:最佳PECS为0.292。瓶颈在于表达而非感知:线性探针在各物理领域从隐藏状态解码可回答性的AUROC最高达0.91;引导单层无效方向可因果性地诱导或抑制弃权。我们的结果在三个开放权重系列(Qwen、Gemma、LLaVA)中均得到复现。此外,该失败在视觉不确定性中比文本不确定性中更显著:模型检测文本不可能的概率约为检测缺失视觉证据的4倍。缩小这种表征-输出差距可能需要输出阶段干预。
英文摘要
When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce TRAPSBench, a procedurally generated video benchmark of 1,404 matched physics pairs in which a single targeted change renders the outcome undeterminable from the visual evidence. Furthermore, we introduce Penalized Epistemic Calibration Score (PECS), a new robust metric that requires models to both answer correctly when the outcome is knowable, and abstain when the outcome is not. Across 16 VLMs spanning five families, spontaneous restraint is poor: the best PECS is 0.292. The bottleneck is expression, not perception: linear probes decode answerability from hidden states at up to 0.91 AUROC across physics domains; steering a single-layer void direction causally induces or suppresses abstention. Our results replicate across three open-weight families (Qwen, Gemma, LLaVA). The failure is also more pronounced in visual than textual uncertainty: models detect textual impossibility about 4x more readily than missing visual evidence. Closing this representation--output gap likely requires output-stage interventions.
发表机构
- Meta Superintelligence Labs(元超级智能实验室)
- Reflection AI(反射人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。