FADE:从被动验证到反事实视频理解中的主动发现
FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding
浏览论文内容
中文总结 AI 辅助
本文提出FADE框架,解决现有反事实视频理解基准泄露目标事件的问题,通过两阶段训练实现主动发现,在三项任务上性能优于GPT-5.6,为相关研究提供坚实基线。
中文摘要 AI 辅助
反事实视频理解用于评估模型是否掌握物理与常识规律。然而,现有多项选择(MCQ)基准会通过问题和候选选项无意中泄露目标事件,这将核心挑战从主动发现降低为文本引导的验证。本文提出FADE,一种用于反事实发现与解释的有效训练框架,其基于证据优先的两阶段训练范式:第一阶段,证据内化的监督微调将模型预测锚定在决定性视觉异常上;第二阶段,应用渐隐锚强化学习策略,逐步移除文本引导,迫使模型独立发现并解释证据。为严格评估该能力,本文还提出一种有效流程,无需额外数据整理即可将现有MCQ数据集转换为对齐的MCQ、开放式问答(OQA)及字幕生成任务。该方法效果显著:以Qwen3-VL-8B为基线,FADE在DualityVidQA-test和IPV-Bench的三项任务上均取得最先进的严格配对分数,性能优于GPT-5.6;从受约束的MCQ转向无约束的OQA和字幕生成时,模型表现出极强鲁棒性,在DualityVidQA-test上的性能保留率分别为90.4%和67.4%,远高于GPT-5.6的48.1%和30.7%。本文希望该简单框架能成为未来无约束反事实视频理解研究的坚实基线。
英文摘要
Counterfactual video understanding evaluates whether models grasp physical and commonsense regularities. However, existing multiple-choice question (MCQ) benchmarks inadvertently leak target events through their questions and candidate options. This reduces the core challenge from active discovery to text-guided verification. In this paper, we present FADE, an effective training framework for counterfactual discovery and explanation. Our method is built on an evidence-first, two-stage training paradigm. First, evidence-internalized supervised fine-tuning grounds the model's predictions in decisive visual anomalies. Second, we apply a fading-anchor reinforcement learning strategy that progressively removes textual guidance, compelling the model to independently discover and explain evidence. To rigorously evaluate this capability, we also introduce an effective pipeline that converts existing MCQ datasets into aligned MCQ, open-ended question answering (OQA), and captioning tasks without requiring additional data curation. Our simple approach yields strong results. Using Qwen3-VL-8B as the baseline, FADE achieves state-of-the-art strict paired scores across all three tasks on DualityVidQA-test and IPV-Bench, outperforming GPT-5.6. In specific, when transitioning from constrained MCQs to unconstrained OQA and captioning, our model demonstrates remarkable robustness. Its performance retention is 90.4% and 67.4% on DualityVidQA-test-substantially higher than the 48.1% and 30.7% retained by GPT-5.6. We hope this simple framework can serve as a solid baseline for future research in unconstrained counterfactual video understanding.