发表机构
University of Houston; The University of Texas at Arlington(休斯顿大学; 德克萨斯大学阿灵顿分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对通用大模型令牌压缩的偏向音频分配问题,提出无需训练的Macer压缩器,先分配音视频预算再按模态排序,降本同时在多基准上保持性能,在OmniVinci-9B上提升达12.9个点。
AI 中文摘要
通用大模型(OmniLLMs)中的令牌压缩通常被表述为单一的显著性排序问题:为每个多模态令牌打分,保留前K个。我们认为这种抽象定义不当。相同的注意力分数同时决定两件事:每个模态获得多少保留容量,以及模态内保留哪些令牌。因此,共享的前K规则继承了偏向音频的分配先验,在视频令牌有竞争机会之前,就将保留容量分配给音频。我们提出Macer,一种无需训练的压缩器,它首先为音频和视频分配明确的预算,然后在特定模态的浅层内执行分配归一化排序。Macer显著降低了令牌成本,同时在音频基础、音视频联合、视觉主导和视频中心基准上保持了准确性。在25%的保留率下,Macer在Qwen2.5-Omni-7B上保留了全令牌性能的98.7%,在Qwen2.5-Omni-3B上保留了97.3%。在Qwen2.5-Omni-7B上,该25%设置在45%保留率时达到了OmniZip级别的性能,同时使用更低的浮点运算量(FLOPs)。在OmniVinci-9B上,相同的“先分配后排序”原则比共享前K排序提升了多达12.9个点。
英文摘要
Token compression in OmniLLMs is typically posed as a single saliency-ranking problem: score each multimodal token, keep the top-K. We argue this abstraction is mis-specified. The same attention score simultaneously decides two things: how much retained capacity each modality receives, and which tokens within a modality are kept. A shared top-K rule therefore inherits this audio-favoring allocation prior, spending retained capacity on audio before video tokens have a chance to compete. We propose Macer, a training-free compressor that first assigns explicit audio and video budgets, then performs allocation-normalized ranking within each modality at modality-specific shallow layers. Macer significantly reduces token cost while preserving accuracy across audio-grounded, audio--video joint, visual-dominant, and video-centric benchmarks. At 25 % retention, Macer preserves 98.7 % of full-token performance on Qwen2.5-Omni-7B and 97.3 % on Qwen2.5-Omni-3B. On Qwen2.5-Omni-7B, this 25 % setting reaches OmniZip-level performance at 45 % retention while using lower FLOPs. On OmniVinci-9B, the same allocation-before-ranking principle improves over shared top-K ranking by up to 12.9 points.