发表机构
Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对全模态大语言模型推理时音频-视频令牌序列长、预填充延迟高和GPU内存使用量大的问题,提出无需训练的查询感知Omni-Prune框架,联合去除冗余并保留跨模态证据,实验证明其性能优于基线方法,能加速预填充并减少内存。
AI 中文摘要
全模态大语言模型正在迅速扩展多模态推理以涵盖同步音频和视频。然而,由此产生的音频-视频令牌序列很长,导致推理时的预填充延迟和GPU内存使用量很高。现有的令牌剪枝方法主要针对仅视觉输入设计,忽略了音频和视频之间的跨模态链接以及决定哪些内容重要的用户查询。为了弥合这一差距,我们提出了Omni-Prune,这是一个无需训练、查询感知的视听令牌剪枝框架,它在保留与任务相关的跨模态证据的同时,联合去除两种模态中的冗余。具体来说,Omni-Prune首先将令牌序列分割成放置在音频显著性峰值处的自适应时间窗口,然后在一个结合编码器注意力和文本查询相关性的单一尺度上对音频和视频令牌进行评分,并对相关的音频-视频令牌进行配对,以便它们保持在一起。在每个窗口内,最后的K-medoids步骤然后选择一些有代表性的令牌,添加基于分数的选择单独会错过的不同线索。广泛的实验表明,Omni-Prune优于既定的基线方法,在保留超过99%的全模型性能的同时,实现了高达3.25倍的预填充加速和1.3倍的内存减少。
英文摘要
Omnimodal large language models (OmniLLMs) are rapidly extending multimodal reasoning to cover synchronized audio and video. However, the resulting audio-video token sequences are long, leading to high prefill latency and GPU memory usage at inference time. Existing token pruning methods, designed mainly for vision-only inputs, miss both the cross-modal links between audio and video and the user query that decides which content matters. To bridge this gap, we present Omni-Prune, a training-free, query-aware audio-visual token pruning framework that jointly removes redundancy from both modalities while keeping task-relevant cross-modal evidence. Specifically, Omni-Prune first splits the token sequence into adaptive time windows placed at audio saliency peaks, then scores audio and video tokens on a single scale that combines encoder attention with text-query relevance, and pairs related audio-video tokens so that they are kept together. Within each window, a final K-medoids step then selects a few representative tokens, adding diverse cues that score-based selection alone would miss. Extensive experiments demonstrate that Omni-Prune outperforms established baseline methods, delivering up to 3.25x prefill speedup and 1.3x memory reduction while retaining over 99% of full-model performance.
Comments14 pages, 7 figures. Code: https://github.com/kimberlyii/Omni-Prune