AI 中文总结
该研究提出无需训练的MEDR帧选择方法,通过多信号事件建模与动态重评分实现查询无关,在Video-MME等基准上提升了多模态大语言模型的长视频处理准确率,且帧集可跨问题复用。
AI 中文摘要
帧选择是多模态大语言模型的核心组件,可在有限的视觉令牌和计算预算下处理长视频。均匀采样能保留时间覆盖范围,但可能遗漏仅短暂出现的信息内容。为缓解这一局限,依赖查询的方法可检索与问题相关的帧。然而,由于所选帧依赖当前问题,相同的视觉输入无法在不同问题间直接共享,多轮视频对话中必须重复进行帧选择。这促使我们研究查询无关的帧选择方法,该方法在保留固定视觉输入可复用性的同时,比均匀采样更好地覆盖信息事件。我们提出MEDR(Multi-Signal Event Modeling and Dynamic Rescoring),一种无需训练且与查询无关的帧选择方法。多信号事件建模将互补的视觉、运动和文本信号组织为信号特定的时间事件;动态重评分则迭代地根据当前所选集合重新评估每个候选,依据帧级信号强度、额外事件覆盖范围和时间邻近性更新其得分。最终构建的固定帧集无需观察查询即可生成,可在不同问题间复用。在标准基准评估中,MEDR在Video-MME上将模型准确率提升0.63%-0.89%;在LongVideoBench的长视频子集上,使用Qwen3-VL-8B时准确率提升最高达1.23%;MEDR对每个视频的所有问题复用完全相同的帧集,同时将整体准确率提升0.53%。
英文摘要
Frame selection is a fundamental component of multimodal large language models, enabling long videos to be processed under limited visual-token and computational budgets. Uniform sampling preserves temporal coverage but may miss informative content that appears only briefly. To alleviate this limitation, query-dependent methods can retrieve question-relevant frames. However, because the selected frames depend on the current question, the same visual input cannot be directly shared across different questions, and frame selection must be repeated in multi-turn video dialogue. This motivates us to seek a query-independent frame selection method that preserves the reusability of a fixed visual input while improving the coverage of informative events beyond uniform sampling. We propose Multi-Signal Event Modeling and Dynamic Rescoring (MEDR), a training-free and query-independent frame selection method. Multi-Signal Event Modeling organizes complementary visual, motion, and text signals into signal-specific temporal events. Dynamic Rescoring then iteratively reevaluates each candidate relative to the current selected set, updating its score according to frame-level signal strength, additional event coverage, and temporal proximity. The resulting fixed frame set is constructed without observing the query and can be reused across different questions. On the standard benchmark evaluations, MEDR improves model accuracy by 0.63%-0.89% on Video-MME. On the long-video subset of LongVideoBench, it improves accuracy by up to 1.23% with Qwen3-VL-8B. MEDR further improves overall accuracy by 0.53%, while reusing exactly the same frame set for every question about a video.
Comments9 pages, 2 figures, 3 tables