arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34772cs.AI

在令牌提交之前:扩散视觉语言模型中视觉幻觉的轨迹级基准测试

Before the Token Commits: Trajectory-Level Benchmarking of Visual Hallucinations in Diffusion VLMs

Yadong Wang, Siping Yue, Yu Tian, Chuanxing Geng, Xiang Chen

首次发表
浏览论文内容

中文总结 AI 辅助

提出轨迹级基准DynaHall和承诺感知协议,发现扩散视觉语言模型的幻觉在令牌承诺前已稳定,并利用PGS方法在掩码状态进行干预以减少假阳性,强调沿生成轨迹而非仅最终输出进行评测与缓解。

中文摘要 AI 辅助

多模态扩散语言模型通过迭代去掩码令牌来生成响应,使得每个答案成为多步轨迹的终点,而非即时承诺。为自回归模型构建的幻觉基准仅评估最终输出,因此无法确定扩散视觉语言模型中的无根据主张是出现较晚,还是在任何答案令牌揭示之前已经稳定。我们引入了DynaHall,一个轨迹级的基准测试,包含基于注释的二元视觉命题,涵盖对象存在性、计数、属性和关系,并通过视觉先验分级的受控硬负样本。DynaHall配备了一个承诺感知协议,记录每个去掩码步骤的中间答案倾向以及最终承诺输出。在来自三个架构族的五个扩散视觉语言模型上,视觉幻觉在承诺之前已经确定:当答案位置仍被掩码时,无根据的答案已是首选状态,后续去掩码步骤很少逆转它,因此失败并非在写入步骤引入。这一现象在解码调度、答案格式和开放式生成中均成立。DynaHall还暴露了最终输出指标隐藏的失败,包括计数和关系崩溃、先验驱动的假阳性,以及方向随类型变化的属性错误。在此诊断的指导下,PGS(承诺前梯度引导)编辑仍被掩码的答案状态以减少假阳性,使肯定率接近平衡,并迁移到另一个架构而不降低通用能力。DynaHall和PGS表明,幻觉应在扩散视觉语言模型的生成轨迹上进行测量和缓解,而不仅仅在最终答案处。

英文摘要

Multimodal diffusion language models generate responses by iteratively unmasking tokens, making each answer the endpoint of a multi-step trajectory rather than an immediate commitment. Hallucination benchmarks built for autoregressive models evaluate only the final output, and therefore cannot determine whether an unsupported claim in diffusion VLMs appears late or has already stabilized before any answer token is revealed. We introduce DynaHall, a trajectory-level benchmark of annotation-backed binary visual propositions covering object existence, counting, attributes, and relations, with controlled hard negatives graded by visual prior. DynaHall is paired with a commitment-aware protocol that records the intermediate answer tendency at every unmasking step alongside the committed output. Across five diffusion VLMs from three architecture families, visual hallucination is settled before commitment: an unsupported answer is already the preferred state while the answer position is still masked, and later unmasking steps rarely reverse it, so the failure is not introduced at the write step. This holds across decoding schedules, answer formats, and open-ended generation. DynaHall also exposes failures hidden by final-output metrics, including counting and relation collapse, prior-driven false positives, and attribute errors whose direction changes by type. Guided by this diagnosis, PGS (Pre-commitment Gradient Steering) edits still-masked answer states to reduce false positives, bringing the affirmation rate close to balance, and transfers to another architecture without degrading general ability. DynaHall and PGS suggest that hallucination should be measured and mitigated along the generation trajectory of diffusion VLMs, not only at the final answer.

发表机构

  • Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
  • State Key Laboratory of Ocean Sensing, ZJU-Hangzhou Global Scientific and Technological Innovation Center, Zhejiang University(浙江大学杭州国际科创中心海洋感知全国重点实验室)
  • Security Capability Center(360安全能力中心)

机构由 AI 辅助整理,请以论文原文为准。

↑