arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mandela-Bench:多模态模型记住规范图像而非真正看见它们

Mandela-Bench: Multimodal Models Remember Canonical Images Instead of Seeing Them

Yicheng Bao, Zhenkun Gao, Xiahui Guo, Mingqian Yang, Xueheng Li, Bangwei Liu, Mingang Chen, Lijun Li, Xuhong Wang, Xin Tan

arXiv 2609.32763首次发表:更新:

AI 中文总结

Mandela-Bench通过1,507张规范图像编辑测试多模态模型,发现模型在高达72.7%的响应中依赖记忆而非视觉证据,仅1/36模型达到知识接地检测标准。

AI 中文摘要

历史照片和其他规范图像现在可以通过单条指令无缝编辑,往往不留下可靠的像素级痕迹。在这种情况下,操纵的唯一证据可能是关于图像所描绘内容的事实。现有基准依赖于生成器伪影、图像-标题不一致、视觉不合理性或外部参考,因此并未测试模型是否能够利用自身世界知识来验证已识别的图像。我们引入Mandela-Bench,包含1,507张规范图像的编辑:1,359个纯知识伪造,每个都违背一个可验证的事实,以及148个无锚点对照,保留编辑过程而不引入事实矛盾,连同474张未触碰的原始图像。我们不仅评分模型是否检测到伪造,还评分其解释是否识别出被插入的实体或被违反的事实。在36个多模态模型中,从0.8B参数到前沿规模,我们发现一致的失败模式。当公众人物从熟悉的照片中被移除时,模型在高达72.7%的响应中仍会说出该人物的名字。一些模型在单独展示时能区分替换面孔与原始面孔,但仍判断完整编辑后的照片为真实。提供真实事件和日期并不能改善基于知识的检测,而在裁剪掉可识别的构图后提供相同信息却能改善。即使在明确的验证提示下,36个模型中只有1个在至少一半的伪造图像上达到KGR标准。这些结果表明,失败不能仅用缺失知识或感知不足来解释。相反,它们与识别偏向于验证所记住的规范图像而非观察到的编辑这一现象一致。

英文摘要

Historical photographs and other canonical images can now be edited seamlessly with a single instruction, often leaving no reliable pixel-level trace. In such cases, the only evidence of manipulation may be a fact about what the image depicts. Existing benchmarks instead rely on generator artefacts, image-caption inconsistencies, visual implausibilities, or external references, and therefore do not test whether a model can use its own world knowledge to verify a recognized image. We introduce Mandela-Bench, containing 1,507 edits of canonical images: 1,359 knowledge-only forgeries, each contradicting one verifiable fact, and 148 anchor-free controls that preserve the editing process without introducing a factual contradiction, together with 474 untouched originals. We score not only whether a model detects a forgery, but whether its explanation identifies the inserted entity or the fact being violated. Across 36 multimodal models, from 0.8B parameters to frontier scale, we find a consistent failure mode. When a public figure is removed from a familiar photograph, models still name that person in up to 72.7% of responses. Some models can distinguish the replacement face from the original when shown in isolation, yet still judge the full edited photograph as authentic. Providing the true event and date does not improve knowledge-grounded detection, whereas providing the same information after cropping away the recognizable composition does. Even under explicit verification prompts, only one of the 36 models meets the KGR criterion on at least half of the forged images. These results suggest that the failures cannot be explained by missing knowledge or inadequate perception alone. Instead, they are consistent with recognition biasing verification toward the remembered canonical image rather than the observed edit.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑