arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27408cs.CVcs.AIcs.CL

视觉语言模型中的能力极限实为读出极限

What Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit

  • Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Alfredo F. Frontera Del Valle

AI总结:

本文发现视觉语言模型基准中的能力限制源于答案读出约定,而非模型本身;通过实验证明坐标答案显著降低性能,并影响模型排名,提出区分可读与不可读约定的方法。

AI中文摘要:

视觉语言模型的基准测试以某种约定提供答案选项:字母、颜色名称或像素坐标。该约定被视为中性的。我们发现事实并非如此,基准测试报告的限制可能属于读出环节而非模型本身。在200张COCO照片上,当位置以英文给出时,Qwen3-VL-4B在九个位置中为指定对象选择正确位置的概率为68.5%,而当相同位置以像素坐标给出时,正确率为20.0%。随机水平为11.1%。代价出现在答案选项为坐标时;在问题中给模型提供坐标反而只损失3.5个百分点,且不显著。该差距在4x4网格上、8位而非4位量化下,以及按对象大小、边界距离和类别划分的每个切片中均存在。它还决定了哪个模型获胜。两个在英文名称下打平的模型在一个坐标系中相差39个百分点,在另一个坐标系中相差54个百分点,且方向相反。在颜色任务中,四个能胜任该任务的开放模型中有三个表现出惩罚;在照片上,三个开放模型中有两个如此,Gemini也如此,在可解析答案上相差11.1个百分点(p = 1e-4)。GPT-4o则没有。为了探究模型是否完全读取坐标,我们为每个坐标附加错误名称并记录模型遵循哪个。以色调角书写的颜色选项被遵循的概率低于随机水平;而标准化像素约定被遵循的概率是随机水平的四倍。这区分了模型能使用和不能使用的约定,尽管它未能预测两种未尝试约定的准确性。五个模型还以五种不同方式命名同一色轮,因此固定的答案词汇在不同模型之间也不是中性的。在这项工作中,有五次我们因评分器与模型对答案形态的分歧而将能胜任的模型误判为无能。我们报告了每个案例。它们是这种现象的缩影。

英文摘要:

Benchmarks for vision-language models offer their answer choices in some convention: a letter, a color name, a pixel coordinate. That convention is treated as neutral. We find it is not, and that the limits a benchmark reports can belong to the readout rather than to the model. On 200 COCO photographs, Qwen3-VL-4B picks the correct one of nine locations for a named object 68.5% of the time when the locations are given in English and 20.0% when the same locations are given as pixel coordinates. Chance is 11.1%. The cost arises when the answer options are coordinates; giving the model a coordinate in the question instead costs 3.5 points and is not significant. The gap holds on a 4x4 grid, under 8-bit rather than 4-bit quantization, and in every slice by object size, boundary distance and category. It also decides which model wins. Two models that tie under English names differ by 39 points in one coordinate system and by 54 in the other, in opposite directions. On the color task, three of the four open models capable of the task show the penalty; on photographs, two of three open models do, and so does Gemini, at 11.1 points on parseable answers (p = 1e-4). GPT-4o does not. To ask whether a model reads a coordinate at all, we attach the wrong name to each one and record which the model follows. Color options written as hue angles are followed below chance; a normalized pixel convention is followed at four times chance. This tells apart conventions a model can use from ones it cannot, though it did not predict accuracy on two untried conventions. Five models also name the same color wheel five different ways, so a fixed answer vocabulary is not neutral across models either. Five times during this work we measured a capable model as incapable because our scorer and the model disagreed about what an answer looks like. We report each case. They are the phenomenon in miniature.

补充信息

↑