arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03918cs.CVcs.AI

何时与何地查看:用于高效长视频理解的自适应视觉证据调度

When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

Ke Li, Jiayu Chen, Maoliang Li, Zihao Zheng, Hailong Zou, Hengyi Zhang, Xuanzhe Liu, Xiang Chen

首次发表
浏览论文内容

中文总结 AI 辅助

提出无需训练的EcoFrame框架,通过熵门控预算调度和注意力引导候选提议实现高效视觉证据调度,在多个长视频理解基准上实现精度与效率的更优权衡。

中文摘要 AI 辅助

高效长视频理解需要视觉-语言模型(VLMs)对选定的少量帧(作为稀疏视觉证据)进行推理。现有的基于相关性的方法依赖于固定帧预算和候选池的静态一次性选择,而基于智能体的调度器则通过代价高昂的多轮推理和交互式搜索实现自适应。我们提出EcoFrame,这是一种低开销、无需训练的查询自适应视觉证据调度框架。EcoFrame利用VLM的推理反馈,确定何时增加帧预算以及在何处搜索额外的候选证据。具体而言,熵门控预算调度利用输出不确定性,在当前证据足够时提前停止,否则逐步扩大帧预算;同时,注意力引导的候选提议将帧级注意力转换为时间先验,在信息丰富的区域实现密集局部搜索,而在注意力分散时保持全局覆盖。在Video-MME、LongVideoBench和MLVU上的实验表明,EcoFrame在多个VLM主干上实现了更好的精度-效率权衡。在Qwen2.5-VL上,EcoFrame的平均精度为64.4,超过了BOLT的63.5,同时比AKS和BOLT快1.85倍;与基于智能体的A.I.R.相比,EcoFrame保持了相当的精度,推理速度最高提升13.5倍。代码将在此httpsURL提供。

英文摘要

Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based methods rely on static one-shot selection with fixed frame budgets and candidate pools, while agent-based schedulers achieve adaptivity through costly multi-round reasoning and interactive search. We propose EcoFrame, a training-free framework for low-overhead query-adaptive visual evidence scheduling. EcoFrame leverages the VLM's inference feedback to determine when to increase the frame budget and where to search for additional candidate evidence. Specifically, entropy-gated budget scheduling uses output uncertainty to stop early when the current evidence is sufficient or progressively expand the frame budget otherwise. Meanwhile, attention-guided candidate proposal converts frame-level attention into a temporal prior, enabling dense local search in informative regions while preserving global coverage when attention is diffuse. Experiments on Video-MME, LongVideoBench, and MLVU demonstrate that EcoFrame achieves a better accuracy--efficiency trade-off across multiple VLM backbones. On Qwen2.5-VL, EcoFrame achieves an average accuracy of 64.4, surpassing BOLT at 63.5, while providing a $1.85\times$ speedup over AKS and BOLT. Compared with the agent-based A.I.R., EcoFrame maintains comparable accuracy with up to a $13.5\times$ inference speedup. Code will be available at https://github.com/AK-DREAM/EcoFrame.

↑