arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10195cs.CVcs.LG

更准确,更少人工:视觉模型中的格式塔分组

More Accurate, Less Human: Gestalt Grouping in Vision Models

Sudhanva Manjunath Athreya, Sai Phani Kumar Malladi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究引入行为测试套件,对比45个不同类型视觉模型与人类在四项分组任务上的表现,发现部分封闭模型的感知组织对齐度低于基准准确率,为可视化研究提供了审核模型的可复用标准。

中文摘要 AI 辅助

人类视觉会将所见内容组织为整体:同色点组合成序列、相似标记聚合成类别、形状补全为可识别对象,这些是可视化设计所依托的格式塔操作。视觉模型是否以这种方式组织视觉内容尚未得到系统测试。我们引入一套行为测试套件,针对四项分组任务(标记-颜色奇数选择、颜色序列计数、轮廓识别、对象奇数选择),将模型与先前感知研究中的人类数据进行评分对比。我们将其应用于五个训练家族的45个模型:监督式、自监督式、对比式视觉语言编码器、开放权重视觉语言模型(VLMs)及封闭基础模型。该套件显示,与人类反应的一致性能捕捉到传统性能指标无法区分的感知组织方面,部分封闭模型的对齐度显著低于其基准准确率所显示的水平。因此,利用已发表的感知数据进行评分,为可视化研究提供了可复用的标准,无需开展新的用户研究,即可审核进入可视化流程的模型是否以人类受众的方式组织视觉内容。

英文摘要

Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable objects. These are the Gestalt operations that visualization design builds on. Whether vision models organize visual content this way has not been systematically tested. We introduce a behavioral battery that scores models against human data from prior perception studies on four grouping tasks: mark-color odd-one-out, color-series counting, silhouette recognition, and object odd-one-out. We apply it to 45 models across five training families: supervised, self-supervised, and contrastive vision-language encoders, open-weight VLMs, and closed foundation models. The battery reveals that agreement with human responses captures aspects of perceptual organization that conventional performance metrics fail to distinguish, with several closed models exhibiting substantially lower alignment than their benchmark accuracy would suggest. Scoring against published perception data therefore gives visualization research a reusable yardstick, requiring no new user study, for auditing whether the models now entering visualization pipelines organize what they see the way their human audience does.

发表机构

  • University of Utah(犹他大学)
  • Siemens AG(西门子公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑