发表机构
Xiamen University; South China University of Technology; Peking University(厦门大学; 华南理工大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MoHallBench基准,系统评估视频大语言模型中的运动幻觉,涵盖共现先验、顺序推理和相似混淆三类来源,揭示动作识别与幻觉抵抗的解耦现象。
AI 中文摘要
视频大语言模型(VideoLLMs)在视频理解方面取得了显著进展,但仍存在与视觉证据不一致的幻觉。现有基准主要关注对象幻觉或粗粒度动作感知,忽略了一个关键的视频特定问题:运动幻觉,即模型推断出视频中不存在的人类运动。我们提出MoHallBench,一个用于诊断VideoLLMs中运动幻觉的基准。MoHallBench系统评估了幻觉的三个主要来源:共现先验、顺序推理和相似混淆。它包含11,306个视频片段和40,493个问答对,涵盖二选一、多选和生成式设置。我们进一步引入了一种带有偏差感知指标的双向提问协议,以减少二值评估中的肯定偏差。对十个近期开源VideoLLMs的实验揭示了动作识别与幻觉抵抗之间的明显解耦,即在正向动作识别上表现良好的模型往往在对抗性负例上失败。在所有设置中,顺序推理幻觉最为严重,表明当前模型倾向于从部分运动线索中过度推断预期结果。我们的分析进一步证实,更强的先验和更细粒度的相似性会显著放大幻觉。我们希望MoHallBench能够促进未来对VideoLLMs中运动幻觉的评估和缓解。
英文摘要
Video Large Language Models (VideoLLMs) have shown strong progress in video understanding, yet they still suffer from hallucinations that are inconsistent with visual evidence. Existing benchmarks mainly focus on object hallucination or coarse action perception, leaving a key video-specific problem underexplored: motion hallucination, in which models infer human motions that are absent from the video. We present MoHallBench, a benchmark for diagnosing motion hallucination in VideoLLMs. MoHallBench systematically evaluates three major sources of hallucination: co-occurrence priors, sequential inference, and similarity confusion. It contains 11,306 video clips and 40,493 question-answer pairs, covering binary-choice, multiple-choice, and generative settings. We further introduce a bi-directional questioning protocol with bias-aware metrics to reduce affirmation bias in binary evaluation. Experiments on ten recent open-source VideoLLMs reveal a clear decoupling between action recognition and hallucination resistance, as models that perform well on positive action recognition often fail on adversarial negatives. Among all settings, sequential inference hallucination is the most severe, showing that current models tend to over-infer expected outcomes from partial motion cues. Our analyses further confirm that stronger priors and finer-grained similarity substantially amplify hallucination. We hope MoHallBench can facilitate future evaluation and mitigation of motion hallucination in VideoLLMs.
Comments19 pages, 5 figures