解耦感知与推理以实现抗幻觉的视频理解
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
浏览论文内容
中文总结 AI 辅助
本文提出DPL模型,通过解耦感知与推理以提升视频理解的抗幻觉能力,引入感知奖励和FAE评估器,有效提高训练后性能和数据效率。
中文摘要 AI 辅助
视频大语言模型通过生成中间推理文本来提升对复杂视频的推理能力。然而,可靠的推理依赖于准确的视频感知。在现有方法中,感知证据与推理文本交织在一起,使直接监督感知过程变得困难。我们主张,可靠的监督需要明确将感知证据与推理分开,以便独立验证感知。为直接监督感知,我们提出了解耦感知与逻辑(DPL),将感知表示为固定格式的证据单元,包含时间戳和视觉描述。这种结构化的表示方法使直接提取感知内容成为可能,并简化了视频片段与奖励评估之间的对齐。基于DPL,我们引入了感知奖励,鼓励抗幻觉性和基于感知的推理。一个事实-aware评估器(FAE)提供反幻觉分数,并实现了与GPT-4o相当的幻觉评估性能。此外,我们通过将感知结果和问题输入参考模型来验证推理一致性。实验表明,通过提供可靠的进程奖励,Video-DPL在3B和7B规模上均能持续提升训练后性能,同时实现更高的数据效率。
英文摘要
Video Large Language Models improve reasoning over complex videos by generating intermediate reasoning text. However, reliable reasoning depends on accurate video perception. In existing approaches, perception evidence is intertwined with reasoning text, making it difficult to directly supervise the perception process. We argue that reliable supervision requires explicitly separating perception evidence from reasoning so that perception can be verified independently. To supervise perception directly, we propose Decoupled Perception and Logic (DPL), which represents perception as fixed-format evidence units containing timestamps and visual descriptions. This structured representation enables direct extraction of perception content and simplifies alignment between video segments and reward evaluation. Building on DPL, we introduce a perception reward that encourages both hallucination resistance and perception-based reasoning. An Factual-Aware Evaluator (FAE) provides anti-hallucination scores and achieves hallucination evaluation performance comparable to GPT-4o. In addition, we validate reasoning consistency by feeding perception results and questions into a reference model. Experiments show that, by providing reliable process rewards, Video-DPL consistently improves post-training performance at both 3B and 7B scales, while delivering higher data efficiency.
发表机构
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。