arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单一排序,任意预算:用于长视频理解的套娃式证据到上下文帧选择框架

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding

Wang Chen, Yu Chen, Xiang Wang, Shuai Li, Jinfa Huang, Xiawu Zheng

arXiv 2608.05707首次发表:更新:

AI 中文总结

本文提出MEC帧选择框架,将长视频帧选择建模为套娃式排序问题,构建单一可复用排序适配任意预算,提升了长视频理解的准确率并降低了选择延迟。

AI 中文摘要

帧选择对于将大型多模态模型(LMMs)应用于长视频至关重要,这是因为长视频存在严重的帧冗余问题,且LMM的上下文窗口有限。由于合适的帧预算会随下游LMM、推理需求和延迟约束而变化,实用的选择器应能适配多种预算。然而,现有方法通常为每个预定义预算单独优化一个帧子集:当预算变化时,之前选择的证据会被替换,而非逐步扩充。按固定分数对帧排序可实现不同预算间的前缀复用,但该方法忽略了不同排序位置的不同作用。本文将长视频帧选择建模为套娃式排序问题:构建单一优先级序列,其小前缀集中了查询条件下的证据,而更大的前缀则保留该证据并添加更广泛的时间上下文。高效构建此类排序本身颇具挑战,因为对长视频进行密集采样并评估帧-查询相关性会产生大量开销。因此,我们提出套娃式证据到上下文(Matryoshka Evidence-to-Context, MEC)帧选择,这是一个无需训练的框架,可构建可复用的稀疏视频索引,通过稀疏探测和局部缩放发现候选帧,并贪婪构建位置自适应排序:早期位置强调证据,后期位置则逐步偏向时间覆盖,同时保持视觉多样性。因此,单一排序可被截断为任意目标预算,无需重新运行选择器。在四个基准和六个帧预算下,MEC的平均准确率比均匀采样提升了3.77个百分点,可与强大的最先进选择器相媲美,且端到端选择延迟降低了47.37%-51.19%。

英文摘要

Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budget varies with the downstream LMM, reasoning demands, and latency constraints, a practical selector should serve multiple budgets. However, existing methods typically optimize an isolated frame subset for each predefined budget: when the budget changes, previously selected evidence may be replaced rather than progressively augmented. A fixed-weight ranking allows prefix reuse across budgets but applies the same weighting at every position, overlooking the distinct roles of early and later ranks. We formulate long-video frame selection as a Matryoshka ranking problem: constructing a single priority sequence whose small prefixes concentrate query-conditioned evidence, while progressively larger prefixes preserve this evidence and add broader temporal context. Efficiently constructing such a ranking is itself challenging, as densely sampling long videos and evaluating frame-query relevance incurs substantial overhead. We therefore introduce Matryoshka Evidence-to-Context (MEC) Frame Selection, a training-free framework that builds a reusable sparse video index, discovers candidates through sparse probing and local zooming, and greedily constructs a position-adaptive ranking - early positions emphasize evidence; later positions progressively favor temporal coverage while preserving visual diversity. A single ranking can thus be truncated to any target budget without rerunning the selector. Across four benchmarks and six frame budgets, MEC improves average accuracy over uniform sampling by 3.77 points, matches strong state-of-the-art selectors, and reduces end-to-end selection latency by 47.37-51.19%.

Comments19 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑