arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10181cs.CVcs.CY

人类视觉与计算机视觉

Human versus Computer Vision

Elena Sirotkina

首次发表
浏览论文内容

中文总结 AI 辅助

本研究对比计算机视觉显著性模型与美国3023名成年人的1140万注视点,发现未训练的中心标记优于所有训练模型,且模型存在人口统计偏差,提出了评估此类系统的标准。

中文摘要 AI 辅助

计算机视觉显著性模型会预测人眼将注视的位置,每张图像对应一张显著性图,价值数十亿美元的预测注意力产业正销售这些图以替代对真实观看者的测量。我针对受众端的领先模型,与来自3023名按全国配额招募的美国成年人的1140万网络摄像头注视点进行对比,这些成年人在观看流传的新闻照片。我发现,未经训练的中心标记优于所有训练过的网络,因为网络在中心之上添加的内容落在这些受众从未注视的位置。剩余的准确性存在系统性偏差,更偏向年轻、白人及温和派观看者,而非年长、黑人及意识形态极端的观看者。我提出了一种前进方向,并基于群体自身的注视揭示模型能否学习该群体,我将其应用于该样本支持的每一个人口统计轴。最终,我展示了决定人们看到什么的系统如何学会平等对待所有人,本研究提供了判断此类主张应遵循的标准。

英文摘要

Computer vision saliency models predict where people will look, one map per image, and a billion-dollar predicted-attention industry sells those maps in place of measuring real viewers. I test the leading models from the audience side, against 11.4 million webcam gaze points from 3,023 US adults recruited to national quotas, viewing circulating news photographs. I show that an untrained central marker outperforms every trained network, because the content the networks add on top of the center falls where these audiences never look. What accuracy remains is systematically biased, favoring younger, White, and moderate viewers over older, Black, and ideologically extreme ones. I propose a way forward and build on what a group's own gaze reveals about whether a model can learn that group at all, and I apply it across every demographic axis this sample supports. Ultimately, I show how systems that decide what people see can learn to see everyone, and this study supplies the standard by which such a claim should be judged.

↑