AI 中文总结
本文提出基于目标MLLM内部注意力证据的动态视觉选择框架EviSelect,通过GRPO优化的随机策略实现高效长视频理解,在三个基准上性能更优,视觉token减少约50%、端到端加速3.9倍。
AI 中文摘要
近期基于多模态大语言模型(MLLM)的长视频理解进展,通过选择与查询相关的帧,缓解了推理时的计算成本及上下文长度受限问题。然而,现有方法主要依赖外部代理评分器与僵化的启发式规则,不可避免地与目标MLLM的内在证据不匹配,且无法适配非均匀的时空信息密度。本文提出一种名为EviSelect的细粒度动态视觉选择框架,其基于目标MLLM的内部注意力证据。我们通过稀疏预填充高效探测视觉证据,作为结构化先验以指导分布感知的动态采样。具体而言,我们利用高度压缩的视觉输入与稀疏注意力,高效近似目标MLLM的注意力图,且与完整对应版本高度对齐。基于该先验衍生的三个互补注意力分量,我们设计了一个轻量选择器,不仅能精确定位与查询相关的时间戳,还能自适应调整局部采样率与空间分辨率。为实现证据条件下的时空采样,我们将选择器建模为随机策略,并通过组相对策略优化(GRPO)在联合准确率-效率奖励下对其进行优化。通过组相对比较,对更低视觉成本下的正确预测给予奖励,我们的方法鼓励策略根据每个视频的信息密度动态分配计算资源。在三个长视频理解基准上,EviSelect相较于现有方法实现了更优性能,同时减少约50%的所选视觉token,并达到3.9倍的端到端加速。
英文摘要
Recent advancements in MLLM-based long-form video understanding have mitigated inference-time computational cost and limited context lengths by selecting query-relevant frames. However, existing approaches predominantly rely on external proxy scorers and rigid heuristic rules, inevitably suffering from misalignment with the target MLLM's intrinsic evidence and failing to accommodate the non-uniform spatiotemporal information density. In this paper, we propose a fine-grained dynamic visual selection framework named EviSelect, grounded in the target MLLM internal attention evidence. Our method efficiently probes visual evidence via sparse prefilling as a structured prior to guide distribution-aware dynamic sampling. Specifically, we efficiently approximate attention maps of the target MLLM using highly compressed visual inputs and sparse attention, well-aligned to the full counterpart. Conditioned on three complementary attention components derived from this prior, we design a lightweight selector that not only precisely locates query-relevant timestamps but also adaptively adjusts the local sampling rate and spatial resolution. To enable evidence-conditioned spatiotemporal sampling, we formulate the selector as a stochastic policy and optimize it via GRPO under a joint accuracy--efficiency reward. By rewarding correct predictions under lower visual cost through group-relative comparisons, our method encourages the policy to allocate computation dynamically according to the information density of each video. Across three long video understanding benchmarks, EviSelect achieves superior performance compared to existing methods while reducing selected visual tokens by about 50\% and achieving a 3.9x end-to-end speedup.
CommentsProject Page: https://zhangbo135.github.io/EviSelect/