arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23193cs.CV

OmniScope:用于全模态大语言模型的模态解耦令牌压缩

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

  • Media Analytics and Computing Lab, Xiamen University(厦门大学媒体分析与计算实验室)
  • Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究所)
  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

Jinsen Su, Yongdong Luo, Yuexiao Ma, Yibo Hu, Meiguang Jin, Xiawu Zheng

AI总结:

研究针对全模态大语言模型令牌压缩问题,提出无需训练的OmniScope框架,以查询为语义锚点分别估计音频和视频相关性,分配特定模态预算并采用相应策略处理,在多基准和模型规模上效果佳,实现加速和内存减少,提出跨模态推理设计原则。

AI中文摘要:

现有的全模态大语言模型令牌压缩方法通常依赖一种模态来决定在另一种模态中保留什么。研究表明这种假设往往不成立,因为相同查询下音频和视频相关性峰值时刻不同,这种跨模态显著性不匹配会导致激进压缩下丢弃关键线索。提出OmniScope,一个无需训练的令牌压缩框架,以查询为共享语义锚点,分别估计音频和视频相关性。它分配特定模态令牌预算,用保留全局上下文和时间变化的锚点-增量策略修剪视觉令牌,合并音频令牌以减少冗余并保持时间连续性。在四个音频-视频基准和两个Qwen2.5-Omni模型规模上,OmniScope在所有压缩设置中实现最佳平均准确率。在总体令牌保留率为25%时,实现高达3.53倍的预填充加速和超过15%的GPU内存减少,平均准确率仅下降0.35个点。这些结果为全模态大语言模型推理提出了一个简单设计原则:跨模态共享查询,但不共享显著性估计。代码可在指定网址获取。

英文摘要:

Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often breaks down: for the same query, audio and video relevance often peaks at different moments. This cross-modal salience mismatch makes unidirectional guidance prone to discarding answer-critical cues under aggressive compression. We propose OmniScope, a training-free token compression framework that uses the query as a shared semantic anchor while estimating relevance separately for audio and video. OmniScope allocates modality-specific token budgets, prunes visual tokens with an anchor-delta strategy that preserves both global context and temporal changes, and merges audio tokens within each second to reduce redundancy while maintaining temporal continuity. Across four audio-video benchmarks and two Qwen2.5-Omni model scales, OmniScope achieves the best average accuracy across all compression settings. At 25% overall token retention, it delivers up to 3.53x prefill speedup and more than 15% GPU memory reduction, with only a 0.35-point drop in average accuracy. These results suggest a simple design principle for OmniLLM inference: share the query across modalities, but not the salience estimates. The code is available at https://github.com/MAC-AutoML/OmniScope.

↑