发表机构
University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Omni2LoRA是一种保持一致性的参数化存储器压缩框架,通过优化秩分配策略提升全模态语言模型效率,在视听问答任务中优于多种基线,大幅缩短推理时间。
AI 中文摘要
全模态语言模型(OLM)可实现统一的视听理解,但处理长联合token序列会使推理的计算成本过高。尽管近期的token压缩方法试图减轻这一负担,但孤立地压缩模态往往会破坏连贯推理所需的时间跨模态锚点。我们提出Omni2LoRA,这是一种通过保持一致性的上下文蒸馏实现高效参数化存储器压缩的两阶段框架,完全绕过了token瓶颈。首先,Perceiver超网络处理冻结OLM的中间表示,在单次前向传播中将多模态上下文编码为满秩低秩适配(LoRA)适配器。为防止所得参数占用量随记录长度线性缩放,我们通过分组相对策略优化(GRPO)优化离散秩分配策略,该策略使用模态消融的反事实奖励明确惩罚视听一致性的损失,迫使模型将其固定的亚线性秩预算分配给协同跨模态锚点,而非孤立的视觉特征。在三个全模态骨干上,采用30%秩预算的Omni2LoRA在四个视听问答基准上的表现优于直接全上下文推理和强大的token压缩基线(OmniZip、OMAC、O-MARC),比最强基线的平均准确率提高8-12%,且在压缩率高达75%时仍保持稳定,而token剪枝方法在此压缩率下会急剧下降。通过将多模态存储器转换为固定预算、可重复使用的参数状态,我们的方法将回答时的多模态token负载降至零,与全上下文推理相比,将每查询的首token时间(TTFT)最多缩短12倍,且在少数查询后摊销至0.5秒以内。
英文摘要
Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally prohibitive. While recent token compression methods attempt to alleviate this burden, compressing modalities in isolation often destroys the temporal cross-modal anchors necessary for coherent reasoning. We introduce Omni2LoRA, a two-stage framework for efficient parametric memory compression via coherence-preserving context distillation that bypasses the token bottleneck entirely. First, a Perceiver hypernetwork processes intermediate representations from a frozen OLM to encode the multimodal context into a full-rank Low-Rank Adaptation (LoRA) adapter in a single forward pass. To prevent the resulting parameter footprint from scaling linearly with recording length, we optimize a discrete rank allocation policy via Group Relative Policy Optimization (GRPO) that uses a modality-ablated counterfactual reward to explicitly penalize the loss of audio-visual coherence, forcing the model to allocate its fixed sub-linear rank budget to synergistic cross-modal anchors rather than isolated visual features. Across three omnimodal backbones, Omni2LoRA operating at a 30% rank budget outperforms direct full-context inference and strong token-compression baselines (OmniZip, OMAC, O-MARC) on four audio-visual question answering benchmarks, improving average accuracy by 8-12% over the strongest baseline and remaining stable under compression ratios as tight as 75%, where token-pruning methods degrade sharply. By converting multimodal memory into a fixed-budget, reusable parameter state, our method drives answer-time multimodal-token load to zero, cutting per-query Time to First Token (TTFT) by up to 12x relative to full-context inference and amortizing to under 0.5s after a handful of queries.