OmniPack:面向高效全模态大语言模型的统一令牌压缩方法
OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models
AI总结:
本文针对全模态大语言模型的高计算开销问题,提出无需训练的OmniPack框架,通过LLM前结构压缩与LLM内语义优化的协作,在5个基准上实现最优性能-效率权衡,在Qwen2.5-Omni-7B上大幅降低计算量且保持高性能。
AI中文摘要:
全模态大语言模型(Omni-LLMs)在视听理解任务上取得了显著性能,但处理长且冗余度高的视听令牌序列会产生大量计算开销,需采用激进的令牌压缩以实现高效部署。现有方法在低令牌预算下常出现性能下降:LLM前的压缩可能丢弃结构重要且全局分布的证据,而LLM内部的压缩往往未能充分利用查询条件下的视听协作。为解决这些局限,本文提出OmniPack,这是一种无需训练的框架,在LLM前协调结构压缩,在LLM内进行任务相关的语义优化。在LLM前,OmniPack通过模态特定重要性、全局覆盖范围及相似度感知合并去除结构冗余;在充分的多模态交互后,它通过文本引导和视听协作进一步整合多样的、任务相关的表示。在包含3种Omni-LLM骨干的5个基准上开展的大量实验表明,OmniPack在不同保留率下均实现了最优的性能-效率权衡,优于所有现有方法。值得注意的是,在Qwen2.5-Omni-7B上,OmniPack在保留98.0%原始性能的同时将FLOPs降至16.7%,且仅用6.8%原始FLOPs仍保留92.9%的原始性能。
英文摘要:
Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhead, demanding aggressive token compression for efficient deployment. Existing methods often degrade at low token budgets: pre-LLM compression may discard structurally important and globally distributed evidence, whereas inner-LLM compression often underexploits query-conditioned audio-visual collaboration. To address these limitations, we propose OmniPack, a training-free framework that coordinates structural compression before the LLM with task-relevant semantic refinement within the LLM. Before the LLM, OmniPack removes structural redundancy through modality-specific importance, global coverage, and similarity-aware merging. After sufficient multimodal interaction, it further consolidates diverse, task-relevant representations through textual guidance and audio-visual collaboration. Extensive experiments on five benchmarks with three Omni-LLM backbones demonstrate that OmniPack consistently achieves the best performance-efficiency trade-off across diverse retention ratios, outperforming all existing methods. Notably, on Qwen2.5-Omni-7B, OmniPack preserves 98.0% of the original performance while reducing FLOPs to 16.7%, and still retains 92.9% of the original performance with only 6.8% of the original FLOPs.