arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

前沿视觉语言模型在检测AI生成肖像方面已超过年轻人,但在校准能力上并未超越

Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibration

Sunwhi Kim, Sunyul Kim, Meounggun Jo, Jini Tae

arXiv 2608.30210首次发表:更新:

发表机构

Hwasung Medi-Science University; Yonsei University; Hoseo University; Gwangju Institute of Science and Technology (GIST)(华城医科大学; 延世大学; 湖西大学; 光州科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究测试19个视觉语言模型检测AI生成肖像的能力,发现2026年7月发布的模型检测准确率已超年轻人,但在校准能力上仍不及人类。

AI 中文摘要

AI图像生成器如今能创建出与真实照片难以区分的人脸肖像,视觉语言模型(VLMs)正被越来越多地用于标记此类图像。我们在相同的198张人脸肖像(真实照片及身份匹配的ChatGPT-4o、Imagen 3版本)上,以与此前针对1667名成年人的研究相同的任务,对19个VLMs进行了基准测试(整体准确率为85%,准确率随年龄增长急剧下降)。2026年6月的14个模型仅达到20至30岁成年人的水平;四周后这一上限被打破,在相同协议下的2026年7月发布的5个模型中,gpt-5.6-sol达到92.8%的平衡准确率(5次抽取均值为92.1%),明显高于20岁成年人的88.5%,claude-fable-5则检测出所有AI图像,平均准确率为91.9%。模型的灵敏度现已远超年轻人(d'最高达3.4,而人类约为2.4),但未被超越的是人类的校准能力:模型的标准范围从c=-1.10到+1.45,而各年龄段人类的标准均接近零;两个新的领先模型均存在偏差(分别为+0.44、-0.97),仅有少数排名中等的模型接近人类的平衡状态。更改标记示例仍会翻转约四分之一的答案。目前最优的机器在此处的视觉检测能力已超过年轻人,但未达到人类在怀疑与信任之间的平衡状态。

英文摘要

AI image generators now create face portraits that are hard to tell from real photographs. Vision-language models (VLMs) are increasingly proposed to flag such images. We benchmarked 19 VLMs on the same 198 face portraits -- real photographs and identity-matched ChatGPT-4o and Imagen 3 versions -- under the same task as our earlier study of 1,667 adults (85% correct overall; accuracy fell steeply with age). The June-2026 cohort of 14 models only matched adults in their 20s-30s. Four weeks later the ceiling broke. Among five July-2026 releases under the identical protocol, gpt-5.6-sol reached 92.8% balanced accuracy (five-draw mean 92.1%), clearly above adults in their 20s (88.5%), and claude-fable-5 detected every AI image while averaging 91.9%. Model sensitivity now exceeds young adults decisively (d' up to 3.4 versus ~ 2.4). What has not been overtaken is human calibration. Model criteria spread from c = -1.10 to +1.45 while humans sit near zero at every age; both new leaders are biased (+0.44, -0.97), and only a few mid-ranked models approach the human balance. Changing the labelled examples still flipped about one answer in four. The best machines now out-see young adults here, without matching the human balance between suspicion and trust.

Comments35 pages, 8 figures, 5 supplementary figures. Human reference data reused (not newly collected) from arXiv:2603.24048. Data and code: https://doi.org/10.5281/zenodo.22148304

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑