ASSEMBLE:面向证据支撑的视频推理的原子技能
ASSEMBLE: Atomic Skills for Evidence-Grounded Video Reasoning
浏览论文内容
中文总结 AI 辅助
ASSEMBLE通过原子技能和证据目录,将显式证据接地融入长视频推理,在三个基准上提升答案准确率与接地准确率。
中文摘要 AI 辅助
复杂的视频推理往往依赖于分散在遥远时刻、实体和事件中的证据,然而仅凭正确答案并不能揭示模型是否依赖了视频中正确的部分。我们提出了ASSEMBLE,一个在整个长视频推理过程中使支撑证据明确化的框架。ASSEMBLE将局部观察和跨片段叙述组织成带有时间戳的证据目录,这些目录可追溯到源视频。随后,一个具有接地感知的阅读器组合出针对问题的原子技能,其结构化输出包含明确的证据引用和支持评估。我们使用正确性门控的引用对齐作为直接的接地信号:在教师监督微调后,组相对策略优化(GRPO)联合优化答案正确性和引用对齐。这产生了可检查的中间轨迹,同时保持最终预测与明确的支撑证据相关联。使用一个由235B教师监督的9B阅读器以及共享的预计算证据目录,ASSEMBLE在三个长视频推理基准上实现了59.2%的宏平均答案准确率,而Gemini-2.5-Pro为58.3%,同时将基于重叠的宏平均接地准确率提高了6.7%,并在所有三个基准上均有所提升。消融实验进一步表明,在相同的后训练阅读器和推理预算下,结构化技能推理比自由形式推理提高了接地准确率。总之,这些结果表明,显式证据接地可以直接集成到长视频推理中,而不会牺牲答案准确率。
英文摘要
Complex video reasoning often depends on evidence scattered across distant moments, entities, and events, yet a correct answer alone does not reveal whether a model relied on the right parts of the video. We introduce ASSEMBLE, a framework that makes supporting evidence explicit throughout long-video reasoning. ASSEMBLE organizes local observations and cross-clip narratives into timestamped evidence catalogs traceable to the source video. A grounding-aware reader then composes question-specific atomic skills whose structured outputs contain explicit evidence references and support assessments. We use correctness-gated citation alignment as a direct grounding signal: after teacher-supervised fine-tuning, Group Relative Policy Optimization (GRPO) jointly optimizes answer correctness and citation alignment. This produces inspectable intermediate traces while keeping final predictions linked to explicit supporting evidence. Using a 9B reader supervised by a 235B teacher and shared precomputed evidence catalogs, ASSEMBLE achieves 59.2% macro-averaged answer accuracy across three long-video reasoning benchmarks, compared with 58.3% for Gemini-2.5-Pro, while improving macro-averaged overlap-based Grounded accuracy by 6.7%, with gains on all three benchmarks. Ablations further show that, with the same post-trained reader and inference budget, structured skill inference improves Grounded accuracy over free-form reasoning. Together, these results show that explicit evidence grounding can be integrated directly into long-video reasoning without sacrificing answer accuracy.
发表机构
- University of Maryland, College Park(马里兰大学帕克分校)
- Google(谷歌)
机构由 AI 辅助整理,请以论文原文为准。