发表机构
Shenzhen International Graduate School, Tsinghua University; Pengcheng Laboratory(清华大学深圳国际研究生院; 鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对全模态大语言模型处理长音视频令牌序列的高预填充成本问题,提出无需训练的两阶段压缩框架OmniRoute,通过时间证据引导预算分配和预算约束语义压缩,在四个基准上实现效率与性能的更优权衡。
AI 中文摘要
全模态大语言模型(Omni-LLMs)将音频和视觉流编码为时间交织的令牌序列,以进行多模态推理。然而,处理长音频-视觉令牌序列会产生大量的预填充成本。现有的压缩方法已取得进展,但往往忽视了音频-视觉语义相关性的时间变化。受时间变化和局部连续性的启发,我们提出了OmniRoute,一种无需训练的两阶段压缩框架。首先,时间证据引导的预算分配(TEGB)根据语义相关性和局部内容变化,推导出分块模态偏好和初始主导模态预算。其次,预算约束的语义压缩(BCSC)压缩主导模态,然后根据实际保留比例校准跟随者的保留目标。对于视频,它结合了时空分组与查询引导选择;对于音频,它基于编码器注意力和查询相关性选择令牌,然后在视觉引导下将残余令牌合并到上下文锚点中。在四个代表性基准上的实验表明,与竞争性基线相比,推理效率和性能之间有更好的权衡。代码和界面将发布以促进进一步研究。
英文摘要
Omnimodal large language models (Omni-LLMs) encode audio and visual streams into temporally interleaved token sequences for multimodal reasoning. However, processing long audio-visual token sequences incurs substantial prefill costs. Existing compression methods have made progress, but often overlook temporal changes in audio-visual semantic relevance. Motivated by temporal variation and local continuity, we propose OmniRoute, a training-free, two-stage compression framework. First, Temporal Evidence-Guided Budgeting (TEGB) derives chunk-wise modality preferences and initial leading-modality budgets from semantic relevance and local content variation. Second, Budget-Constrained Semantic Compression (BCSC) compresses the leading modality and then calibrates the follower's retention target using the actual retained fraction. For video, it combines spatiotemporal grouping with query-guided selection; for audio, it selects tokens based on encoder attention and query relevance, then merges residual tokens into context anchors under visual guidance. Experiments on four representative benchmarks demonstrate a better trade-off between inference efficiency and performance than competitive baselines. The code and interface will be released to facilitate further research.