AI 中文总结
研究多模态视频错误信息检测,提出SIEVE框架,通过智能体探索多模态证据构建证据包,由验证器判断真实性。智能体经特殊训练,实验表明该框架性能优于基线,能提供可检查证据线索,提升检测透明度与可信度。
AI 中文摘要
多模态视频错误信息检测通常被视为一个整体视频理解任务,对整个视频及其相关内容一次性处理和判断。然而,现实世界中的错误信息往往呈现稀疏且具有组合性的证据结构,详尽的多模态推理可能会引入大量冗余并模糊决定性证据。这促使将证据获取与验证解耦,先识别稀疏且与决策相关的线索,再基于获取的证据判断真实性。为此,我们提出了SIEVE框架,用于多模态视频错误信息检测中的稀疏交互式证据验证。一个证据搜索智能体积极探索可用的多模态证据并构建一个紧凑的证据包,然后由验证器用于确定真实性。该智能体通过有监督的证据搜索轨迹和一个证据感知强化学习目标进行训练,促进获取信息丰富的证据,同时抑制不必要或无效的交互。在多个视频错误信息基准上的实验表明,SIEVE始终优于评估的基线,并支持使用紧凑证据包进行可靠验证。此外,由此产生的获取过程提供了一个明确且可检查的证据线索,提高了多模态错误信息检测的透明度和可信度。
英文摘要
Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content are processed and judged in a single pass. However, real-world misinformation often exhibits a sparse and compositional evidence structure: a reliable decision may depend on only a few coupled clues, while most video content contributes limited additional information. Exhaustive multimodal reasoning may therefore introduce substantial redundancy and obscure decisive evidence. This motivates decoupling evidence acquisition from verification: first identifying sparse, decision-relevant clues and then judging veracity based on the acquired evidence. Accordingly, we propose SIEVE, a framework for Sparse Interactive Evidence Verification via Extraction in multimodal video misinformation detection. An evidence-seeking agent actively explores the available multimodal evidence and constructs a compact evidence package, which is then used by a verifier to determine veracity. The agent is trained with supervised evidence-seeking trajectories and an evidence-aware reinforcement learning objective that promotes informative evidence acquisition while discouraging unnecessary or invalid interactions. Experiments on multiple video misinformation benchmarks show that SIEVE consistently outperforms the evaluated baselines and supports reliable verification using compact evidence packages. Moreover, the resulting acquisition process provides an explicit and inspectable evidence trail, improving the transparency and groundedness of multimodal misinformation detection.