发表机构
Ant Group; Tsinghua University(蚂蚁集团; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出WorldFact-Bench基准评估图像-世界一致性,提出PERSIST-Agent智能体,经实验证实该智能体可提升对准确率,凸显状态引导验证的价值及视觉合理性与事实正确性的差距。
AI 中文摘要
图像生成技术的进步使得视觉真实性越来越难以评估。尽管图像取证目前会检查生成伪影和更高级别的视觉不一致性,但一张看似合理的图像仍可能与现实世界的事实或规则相矛盾。我们推出WorldFact-Bench,用于在无需预先定义声明或验证目标的情况下,从单张图像评估图像-世界一致性。该基准包含1274个源对齐的真实-虚假对,涵盖四个验证机制和十个语义领域。每个对引入一个特定的、有证据支持的事实冲突,同时力求保留非目标内容和视觉合理性。图像被独立评估,而对准确率要求对一个对的两个成员都进行正确分类。我们进一步提出PERSIST-Agent,该智能体围绕一个持久状态组织迭代验证,该状态连接候选事实、视觉观察、证据和验证状态。此状态指导后续检查和检索,同时保留未解决的候选。在骨干权重固定的情况下,利用自优化通过验证反馈优化智能体的提示和执行规则。实验显示,多个检测器存在强烈的标签偏差,且检索带来的增益不均。在评估的8B骨干模型上,PERSIST-Agent相较于直接判断和检索增强基线,提高了对准确率,而消融实验支持持久验证状态的作用。这些发现凸显了状态引导验证的价值,以及视觉合理性与事实正确性之间仍存在的差距。
英文摘要
Advances in image generation have made visual authenticity increasingly difficult to assess. Although image forensics now examines both generation artifacts and higher-level visual inconsistencies, a plausible image can still contradict real-world facts or rules. We introduce WorldFact-Bench to evaluate image-world consistency from a single image, without a predefined claim or verification target. The benchmark contains 1,274 source-aligned real-fake pairs across four verification regimes and ten semantic domains. Each pair introduces a specific, evidence-supported factual conflict while seeking to preserve non-target content and visual plausibility. Images are evaluated independently, and pair accuracy requires both members of a pair to be classified correctly. We further propose PERSIST-Agent, which organizes iterative verification around a persistent state linking candidate facts, visual observations, evidence, and verification statuses. This state guides subsequent inspection and retrieval while retaining unresolved candidates. With backbone weights fixed, harness self-optimization refines the agent's prompts and execution rules through validation feedback. Experiments reveal strong label biases in several detectors and uneven gains from retrieval. On the evaluated 8B backbones, PERSIST-Agent improves pair accuracy over both direct judgment and retrieval-augmented baselines, while ablations support the role of persistent verification state. These findings highlight the value of state-guided verification and the remaining gap between visual plausibility and factual correctness.