AI 中文总结
针对长视频问答中视觉令牌成本高的问题,提出令牌预算受限的视频证据索引(VEI),通过低分辨率预览定位关键帧并构建高分辨率证据集,结合特权自蒸馏提升有限预算下的问答准确性与证据定位。
AI 中文摘要
长视频问答受到视觉令牌高昂成本和当前视觉语言模型固定上下文宽度的限制。长视频问题可能需要广泛的时间覆盖,但答案往往仅由一小段关键时刻支持。为了高效定位这些时刻,我们提出了令牌预算受限的视频证据索引(VEI):给定一个密集的低分辨率视频预览,模型构建一个紧凑的高分辨率证据集用于最终推理。我们将VEI视为一个策略,必须联合解决证据定位(找到与问题相关的时刻)和预算规划(决定在有限的高分辨率帧预算中何处花费)。我们通过推理流程实现这一想法:视频预览提供廉价的全局覆盖,视频证据索引构建证据集,答案生成结合两者进行最终视觉问答。为解决帧级监督缺失问题,我们采用特权自蒸馏,其中答案感知的教师指导正常测试时策略在由学生生成的索引轨迹上进行。我们探索了每帧1、6、12和24个视觉令牌的预览,训练一个支持所有四种分辨率的单一策略。实验表明,视频证据索引在有限视觉预算下提高了准确性,自蒸馏进一步改善了问答准确性和时间证据定位。
英文摘要
Long-video question answering is limited by the high cost of visual tokens and by the fixed context width of current VLMs. A long-video question may require broad temporal coverage, but the answer is often supported by only a compact set of moments. To locate these moments efficiently, we propose token-budgeted Video Evidence Indexing (VEI): given a dense low-resolution Video Preview, the model constructs a compact high-resolution Evidence Set for final reasoning. We treat VEI as a policy that must jointly solve \textit{evidence localization}, which finds question-relevant moments, and \textit{budget planning}, which decides where to spend the limited high-resolution frame budget. We implement this idea with an inference pipeline: the Video Preview provides cheap global coverage, Video Evidence Indexing constructs the Evidence Set, and Answer Generation combines both inputs for final VQA. To address missing frame-level supervision, we adopt privileged self-distillation, where an answer-aware teacher guides the normal test-time policy on student-generated indexing traces. We explore previews at 1, 6, 12, and 24 visual tokens per frame, training a single policy that supports all four resolutions. Experiments show that Video Evidence Indexing improves accuracy under limited visual budgets, and self-distillation further improves both QA accuracy and temporal evidence localization.