TRACE:用于长视频理解的基于锚定收敛证据的时序检索
TRACE: Temporal Retrieval with Anchored and Convergent Evidence for Long-Horizon Video Understanding
AI总结:
针对长视频理解中证据核查不足的问题,提出VES-Bench基准和无需训练的TRACE智能体,在帧成本更低的情况下实现了高正确率,在多个基准上表现出色。
AI中文摘要:
长视频的答案仅在视频解码出的帧覆盖答案依赖的所有事件时才具备证据支撑性。现有评估仅对最终答案正确性或预测的证据区间打分,但很少核查模型在回答前解码的帧,导致正确答案仍可能基于不完整的观察。我们推出VES-Bench,这是一个包含600个问题的基准,涉及348个公开长视频上的时序排序与事件计数任务,每个问题都带有一组联合必要的证据区间,使我们能在三个严格程度上核查模型解码的帧是否覆盖所有证据区间。我们还提出TRACE,这是一个无需训练的智能体,其答案基于原始视觉片段,逐轮构建证据束,仅当证据束增长时答案稳定且对相同片段的最终核查返回相同答案时才停止。在相同骨干模型的核查下,TRACE在每个证据区间内至少解码2帧时,50.7%的问题回答正确,每问题平均解码98.7帧;相比每问题解码128帧的均匀解码(正确率40.2%,TRACE超出10个百分点以上),每问题解码256帧的均匀解码(正确率63.5%),TRACE的正确率仅低2.6个百分点,但仅为其帧成本的0.39倍,同时在核查中达到最高的答案正确率63.5%。TRACE在Video-MME(86.1)、LVBench(75.6)和LongVideoBench(75.1)上也保持竞争力。
英文摘要:
A long-video answer is evidence-supported only when the frames decoded from the video cover every event the answer depends on. Existing evaluations score final-answer correctness or predicted evidence intervals, but the frames a method decodes before answering are rarely audited, so correct answers can still rest on incomplete observation. We introduce VES-Bench, a 600-question benchmark of Temporal Ordering and Event Counting items over 348 public long videos. Each item carries a jointly necessary set of evidence intervals, letting us audit at three strictness levels whether a method's decoded frames cover every one of them. We also propose TRACE, a training-free agent that grounds answers in raw visual clips, builds an evidence bundle round by round, and stops only when the answer stabilises as the bundle grows and a final pass over the same clips returns the same answer. Under a same-backbone audit, TRACE answers 50.7% of questions correctly with at least two decoded frames inside every evidence interval, at 98.7 frames per question: over 10 points above uniform decoding at 128 frames (40.2%), and within 2.6 points of uniform decoding at 256 frames at 0.39x its frame cost, while reaching the highest answer accuracy in the audit (63.5%). TRACE also stays competitive on Video-MME (86.1), LVBench (75.6), and LongVideoBench (75.1).