State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal Models
具有门控注意力和可学习采样的状态空间分层压缩用于大多模态模型中的小时级视频理解
专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV
AI总结 本文提出了一种基于状态空间模型和门控注意力的高效压缩方法,用于减少大模型中小时级视频的token消耗,同时保持性能。
Comments AAAI 2026 (Oral). Project page: https://github.com/naver-ai/mambamia