arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmniDelta:用于OmniLLMs中令牌压缩的技能驱动预算分配

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

Haoyang Huang, Wenjie Huang, Tianqi Xu, Hongyaoxing Gu, Kang Tan, Yikai Fu, Yuhao Shen, Tianyu Liu, Baolin Zhang, Jun Zhang, Xinyi Hu, Jun Dai, Shuang Ge, Lei Chen, Yue Li, Mingchen Wang, Meng Zhang

arXiv 2607.25669首次发表:更新:

发表机构

Zhejiang University; Carnegie Mellon University; University of Chinese Academy of Sciences; University of Science and Technology of China; Alibaba(浙江大学; 卡内基梅隆大学; 中国科学院大学; 中国科学技术大学; 阿里巴巴)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对OmniLLMs长音频-视频令牌序列成本高的问题,提出技能驱动的OmniDelta框架,通过构建技能池、结合意图与内容感知分配预算,能与现有策略结合,实验证明其在多个基准上建立新的准确性-效率前沿,减少内存并加速推理。

AI 中文摘要

新兴的多模态大语言模型(OmniLLMs)能统一理解文本、音频和视频,但长音频-视频令牌序列带来巨大内存和推理成本。现有压缩方法主要关注固定预算下选择重要令牌,忽视了预算分配问题。我们表明直接查询与音频/视频的相似性对跨模态预算分配不可靠,均匀的模态内预算会遗漏关键证据并保留冗余内容。为解决这些限制,我们提出OmniDelta,一个无需训练、技能驱动的框架,将意图感知的跨模态分配与内容感知的模态内分配相结合。OmniDelta首先构建音频和视频技能池,根据查询需求转移固定保留令牌预算,然后使用局部复杂度和时间冗余在音频段和视频帧上重新分配模态预算。所得局部预算可与现有剪枝策略结合,在保持总保留令牌比例的同时改变预算支出位置。在四个音频-视频基准上对两个Qwen2.5-Omni模型进行的实验表明,OmniDelta在剪枝率上建立了新的准确性-效率帕累托前沿。在Qwen2.5-Omni-7B上保留25%令牌时,OmniDelta将GPU内存减少22.0%,并实现比全令牌推理快1.64倍的端到端加速。

英文摘要

Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly focus on selecting important tokens under fixed budgets, leaving the preceding budget-allocation problem underexplored. We show that direct query-to-audio/video similarity is unreliable for inter-modal budget allocation, and that uniform intra-modal budgets can miss key evidence while retaining redundant content. To address these limitations, we propose OmniDelta, a training-free, skill-driven framework that couples intent-aware inter-modal allocation with content-aware intra-modal allocation. OmniDelta first constructs audio and video skill pools to shift the fixed retained-token budget according to query demand, then reallocates modality budgets over audio segments and video frames using local complexity and temporal redundancy. The resulting local budgets can be combined with existing pruning strategies, preserving the total retained-token ratio while changing where the budget is spent. Experiments on four audio-video benchmarks with two Qwen2.5-Omni models show that OmniDelta establishes a new accuracy-efficiency Pareto frontier across pruning ratios. At 25% token retention on Qwen2.5-Omni-7B, OmniDelta reduces GPU memory by 22.0% and achieves a 1.64x end-to-end speedup over full-token inference.

Comments24 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑