发表机构
School of Electronic and Computer Engineering, Peking University; The Hong Kong University of Science and Technology(北京大学电子与计算机工程学院; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出EFlow框架,通过分离时间定位与逻辑推理(CoT)及置信度感知的反思机制,解决长视频推理中早期语义假设导致的证据偏差问题,在五个基准上提升性能。
AI 中文摘要
长视频推理从根本上受限于模型如何获取和利用视觉证据。现有的工具增强视频框架通常将时间定位和答案推理交织在单一轨迹中,导致早期语义假设偏向证据定位。我们将这种失败模式称为早期语义承诺,即有偏的定位检索到不完整的证据,而不完整的证据进一步强化错误的推理。为了解决这个问题,我们提出了EFlow,一个基于Qwen3-VL构建的证据优先视频推理框架。EFlow通过时间定位的CoT和推理的CoT明确分离时间定位和逻辑推理,使模型能够在答案推理之前检索相关证据。此外,EFlow引入了一种置信度感知的反思机制,当检索到的证据可能不足时重新评估整个视频。我们进一步构建了专门的轨迹数据集,并通过监督微调、强化学习和强化微调训练EFlow。在五个视频理解基准上的大量实验表明,EFlow持续提升了长视频推理性能。
英文摘要
Long-video reasoning is fundamentally constrained by how models acquire and utilize visual evidence. Existing tool-augmented video frameworks often interleave temporal grounding and answer reasoning within a single trajectory, causing early semantic hypotheses to bias evidence localization. We term this failure mode premature semantic commitment, where biased grounding retrieves incomplete evidence and incomplete evidence further reinforces incorrect reasoning. To address this issue, we propose EFlow, an evidence-first video reasoning framework built upon Qwen3-VL. EFlow explicitly separates temporal grounding and logical reasoning through CoT for Temporal Grounding and CoT for Reasoning, enabling the model to retrieve relevant evidence before answer inference. In addition, EFlow introduces a confidence-aware reflection mechanism that re-evaluates the full video when retrieved evidence is potentially insufficient. We further construct dedicated trajectory datasets and train EFlow through supervised fine-tuning, reinforcement learning, and reinforcement fine-tuning. Extensive experiments across five video understanding benchmarks demonstrate that EFlow consistently improves long-video reasoning performance.