arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NumerosityVLM:一种受认知启发的、用于解释视觉-语言模型中数量表征的基准测试

NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models

Yiming Fu, Fangjun Li, Xiujin Liu, Ruidong Ma, Hang Yu, Zhichen Lu, Kanwei He, Alessandro Di Nuovo, Angelo Cangelosi, Zhegong Shangguan

arXiv 2608.15425首次发表:更新:

发表机构

The University of Manchester; University of Michigan; Sheffield Hallam University; Tufts University; ENSTA, Institut Polytechnique de Paris(曼彻斯特大学; 密歇根大学; 谢菲尔德哈勒姆大学; 塔夫茨大学; 巴黎综合理工学院国立高等先进技术学校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对视觉-语言模型数量感知研究的基准缺陷,提出受认知启发的NumerosityVLM基准,评估7个模型发现架构对性能影响最大,数量信号出现在视觉编码器早期,差异主要来自语言模型组件。

AI 中文摘要

视觉-语言模型(VLMs)在高级多模态任务上表现出色,但数量感知(一种人类婴儿在语言习得前就具备的认知能力)在当前模型中仍未得到充分理解,因为现有的计数基准将数量与相关视觉因素混淆了。我们提出了一种受认知启发的诊断基准NumerosityVLM,它包含6种受控条件下的10800张合成图像。该基准正交地操纵物体大小、空间排列和数量,同时逐步去除纹理、形状和颜色。在零样本设置下评估7个VLMs,多因素分析显示,模型架构解释了性能方差的最大比例(偏ω²=0.325),远超视觉条件。逐层探测进一步表明,线性可分的数量信号始终出现在视觉编码器的早期阶段,而被评估模型间的性能差异主要与语言模型组件相关。代码和数据可在this https URL和this https URL获取。

英文摘要

Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before language acquisition, remains poorly understood in current models, as existing counting benchmarks entangle numerosity with correlated visual factors. We introduce a cognitively inspired diagnostic benchmark, NumerosityVLM, comprising 10,800 synthetic images across six controlled conditions. The benchmark orthogonally manipulates object size, spatial arrangement, and numerosity, while progressively ablating texture, shape, and color. Evaluating seven VLMs in a zero-shot setting, multi-factor analysis reveals that model architecture explains the largest proportion of performance variance (partial $ω^{2}=0.325$), far exceeding visual conditions. Layer-wise probing further shows that linearly separable numerosity signals consistently emerge at early stages of the vision encoder, while performance differences across evaluated models are primarily associated with the language model component. Code and data are publicly available at https://github.com/fuy3/NumerosityVLM-Benchmark, and https://huggingface.co/datasets/fuy3/NumerosityVLM.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑