arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12557cs.CV

用于长视频理解中事件感知视觉分配的高斯混合模型

Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding

Yifan Lu, Ziqi Zhang, Chunfeng Yuan, Jun Gao, Bing Li, Weiming Hu

首次发表
浏览论文内容

中文总结 AI 辅助

针对长视频理解中视觉分配问题,提出GMM-EVA方法,利用高斯混合模型建模事件级结构,采用差异化分配策略,在多个长视频基准实验中显著优于均匀采样,以约一半视觉令牌预算达可比性能,凸显高效性。

中文摘要 AI 辅助

大型视觉语言模型在长视频理解中面临挑战,因均匀采样计算成本高且信息损失大。现有关键帧选择方法将视频帧视为原子实体并平均分配视觉预算,忽略高层语义结构并引入冗余。我们提出GMM-EVA,利用高斯混合模型从离散帧观测中建模事件级结构。应用差异化分配策略,为每个事件保留一个高分辨率主关键帧以保留高保真细节,同时利用低分辨率次关键帧维持时间上下文并优化令牌预算。GMM-EVA是无训练、即插即用框架,在各种相关性度量和下游LVLMs上都能稳健泛化。在多个长视频基准上的大量实验表明,我们的方法显著优于均匀采样,在使用约一半视觉令牌预算时性能与基线选择方法相当,突出了其卓越的效率和有效性。

英文摘要

Large Vision-Language Models (LVLMs) face significant challenges in long video understanding due to the excessive computational cost and information loss associated with uniform sampling. Existing keyframe selection methods often treat video frames as atomic entities and allocate visual budgets equally, thereby overlooking high-level semantic structures and introducing substantial redundancy. To address these limitations, we propose GMM-EVA (Gaussian Mixture Modeling for Event-Aware Visual Allocation), which leverages Gaussian Mixture Models to model event-level structure from discrete frame-wise observations. A differentiated allocation strategy is then applied to preserve one primary high-resolution keyframe per event for high-fidelity detail, while utilizing lower-resolution secondary keyframes to maintain temporal context and optimize token budgets. GMM-EVA is a training-free, plug-and-play framework that generalizes robustly across various relevance measures and downstream LVLMs. Extensive experiments on multiple long video benchmarks demonstrate that our method significantly outperforms uniform sampling. Notably, GMM-EVA achieves comparable performance to baseline selection methods while utilizing only approximately half of the visual token budget, highlighting its superior efficiency and effectiveness.

发表机构

  • Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information, CASIA(中国科学院自动化所多模态信息超智能安全北京市重点实验室)
  • State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(中国科学院自动化所多模态人工智能系统国家重点实验室)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
  • Hello Group(未知(保留英文))
  • School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑