MetaSampling:使帧采样器高效用于长视频问答
MetaSampling: Making Frame Samplers Efficient for Long-Video Question Answering
- University of Southern Mississippi(南密西西比大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MetaSampling是一种无需训练的即插即用采样策略,通过动态减少传递给多模态大语言模型的帧数,在36种配置中平均减少8.9%的帧并提升25种配置的准确性,从而提升长视频问答效率。
AI中文摘要:
帧选择是多模态大语言模型(MLLMs)进行长视频问答(VQA)的重要组件。现有的帧选择方法优于简单的top-$k$嵌入检索和均匀采样,但通常在固定的全局选择预算下应用。我们引入了MetaSampling,一种无需训练、即插即用的采样策略,可应用于现有帧选择器之上。MetaSampling通过动态减少传递给MLLM的帧数来提高下游VQA效率,同时保持甚至在某些情况下提高答案准确性。我们在36种帧选择器-MLLM骨干-VQA基准配置组合上评估了MetaSampling。MetaSampling在所有36种配置中减少了所选帧的数量,并在其中25种配置中提高了准确性,平均帧减少8.9%,同时整体准确性略有提升。
英文摘要:
Frame selection is an important component of long-video question answering (VQA) with Multimodal Large Language Models (MLLMs). Existing frame-selection methods improve over simple top-$k$ embedding retrieval and uniform sampling, but are typically applied under a fixed global selection budget. We introduce \textbf{MetaSampling}, a training-free, plug-and-play sampling strategy that can be applied on top of existing frame selectors. MetaSampling improves downstream VQA efficiency by dynamically reducing the number of frames passed to the MLLM while preserving, and in some cases improving, answer accuracy. We evaluate MetaSampling across 36 paired frame-selector--MLLM-backbone--VQA-benchmark configurations. MetaSampling reduces the number of selected frames in all 36 configurations and improves accuracy in 25 of them, yielding an average frame reduction of $8.9\%$ while slightly improving accuracy overall.