基于注意力的MLLM选择器在测试时对长视频进行高效帧选择
Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors
- ZJU(浙江大学)
- NJU(南京大学)
- CSU(中南大学)
- Microsoft(微软公司)
- NJUST(南京理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究利用MLLMs中验证选择提取层的跨模态注意力构建无训练的DAFS帧选择器,通过查询条件聚合提取帧级证据,将候选池大小和每帧令牌预算联合分配问题用动态规划解决,在Video-MME上表现出色,且无需重新训练即可跨多种模型和任务泛化。
AI中文摘要:
使用多模态大语言模型(MLLMs)理解长视频需要从数千个候选帧中选择一组紧凑的帧,但选择正确的帧似乎又需要先理解视频,存在循环依赖。通过观察发现MLLMs中验证选择的提取层的跨模态注意力已提供与查询相关的帧证据,无需自回归生成。利用此特性构建了无训练的帧选择器DAFS。通过查询条件聚合将所选层注意力转换为相关性分数,实现跨帧比较。还将候选池大小和每帧令牌预算的联合分配制定为离散优化问题并通过动态规划解决。在32帧预算下,该选择器在Video-MME上比均匀采样提高了6.4分,优于基于训练的选择器。
英文摘要:
Understanding long videos with multimodal large language models (MLLMs) requires selecting a compact set of frames from thousands of candidates, yet identifying the right frames seemingly requires understanding the video first. We resolve this circular dependency with a simple observation: cross-modal attention at validation-selected extraction layers in MLLMs already provides query-relevant frame evidence without requiring autoregressive generation. We exploit this property to build DAFS (Dynamic Attention-based Budget-aware Frame Selection), a training-free frame selector. A lightweight MLLM selector, even with only 2B parameters, can extract frame-level evidence by converting selected-layer attention into relevance scores through query-conditioned aggregation. This enables cross-frame comparison without autoregressive decoding. To handle the selector's own context constraint, we formulate the joint allocation of candidate pool size and per-frame token budget as a discrete optimization problem solved by dynamic programming. Under a 32-frame budget, our selector improves over uniform sampling by up to 6.4 points on Video-MME and outperforms prior training-based selectors under matched frame budgets, while generalizing across selector and answerer backbones, and across tasks, without retraining.