arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向全模态大语言模型(Omni-LLMs)的基于局部视听动态的延迟音频剪枝方法

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs

Kyeongyoon Lee, Hongyeob Kim, Youngeun Kim, Sungeun Hong

arXiv 2608.08794首次发表:更新:

AI 中文总结

针对全模态大语言模型的高开销问题,提出两阶段框架A-PACK,利用局部视听动态延迟音频剪枝,在Qwen2.5-Omni基准上实现性能最优,同时大幅降低预填充开销并提升解码吞吐量。

AI 中文摘要

全模态大语言模型(Omni-LLMs)可联合处理音频、视频与文本,但长多模态序列会产生巨大的预填充(prefill)与键值缓存(KV-cache)开销。现有全模态压缩方法主要聚焦于大语言模型(LLM)前的 token 减少,而 LLM 边界处的模态特定压缩被研究不足。我们提出 A-PACK,这是一个两阶段框架,它将音频剪枝延迟至查询条件下的多模态交互出现时。我们的分析表明,音频每个 token 相比视频具有更高的任务相关信息密度与表征多样性。我们进一步发现,局部视听动态相比逐 token 匹配能为视觉选择提供更有效的线索。因此我们在 LLM 前保留音频并利用局部动态压缩视频,随后在 LLM 内逐步剪枝低相关性的音频与视觉 token 及其 KV-cache 条目。在 Qwen2.5-Omni-7B/3B 的四个基准测试中,A-PACK 在评估的现有方法中取得了最强的平均性能,同时将预填充 FLOPs 最高降低 78%,将解码吞吐量最高提升 2.21 倍。

英文摘要

Omni-modal LLMs jointly process audio, video, and text, but long multimodal sequences incur substantial prefill and KV-cache costs. Existing omni-modal compression methods primarily focus on pre-LLM token reduction, leaving modality-specific compression across the LLM boundary underexplored. We propose A-PACK, a two-stage framework that defers audio pruning until query-conditioned multimodal interactions emerge. Our analysis shows that audio exhibits higher task-relevant information density and representational diversity per token than video. We further find that local audio-visual dynamics provide a more effective cue for visual selection than token-wise matching. We therefore preserve audio and compress video with local dynamics before the LLM, then progressively prune low-relevance audio and visual tokens and their KV-cache entries inside the LLM. Across four benchmarks on Qwen2.5-Omni-7B/3B, A-PACK achieves the strongest average performance among the evaluated prior methods while reducing prefill FLOPs by up to 78% and improving decoding throughput by up to 2.21x.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑