arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39953cs.CV

学习在压缩上下文中推理:通过自蒸馏实现OmniLLMs的无真值适配

Learning to Reason with Compressed Context: Ground-Truth-Free Adaptation of OmniLLMs via Self-Distillation

Jianghao Wang, Ke Meng, Jian Li, Chi Cheng, Longyu Qi, Liyin Liang, Yifeng Qian, Chunbo Lai, Yutian Lin, Zeyu Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对OmniLLMs压缩上下文推理性能下降问题,提出无真值自蒸馏框架CAFD,利用全上下文自教师监督压缩学生,在多个基准和预算下显著提升准确率-效率权衡。

中文摘要 AI 辅助

全模态大语言模型(OmniLLMs)实现了统一的音视频理解,但其长多模态token序列使得部署计算成本高昂。token压缩降低了这一成本,然而激进的压缩往往会降低准确性。现有工作主要集中于设计更好的压缩机制;然而,适配底层语言模型以在剩余压缩上下文中有效推理仍未被充分探索。为解决此问题,我们提出了CAFD(通过全上下文蒸馏进行压缩上下文适配),一种无真值自蒸馏框架,它将OmniLLMs适配到固定的压缩流水线,无需参考答案、理由或正确性奖励。CAFD利用同一多模态样本的全token视图作为特权信息来源:一个全上下文自教师沿着学生的在线策略轨迹为压缩上下文学生提供软目标监督。在Qwen2.5-Omni-7B上,跨越五个音视频基准、五个压缩流水线和五个部署预算进行评估,CAFD展现出持续的增益,在125个条件中改善了120个,平均准确率提升1.44个百分点,并平均恢复了26.9%的准确率差距。这些结果表明,所提出的无真值适配为提高已部署OmniLLMs的准确率-效率权衡提供了一条有效且实用的途径。

英文摘要

Omni-modal large language models (OmniLLMs) enable unified audio-video understanding, but their long multimodal token sequences make deployment computationally expensive. Token compression reduces this cost, yet aggressive compression often lowers accuracy. Existing works predominantly focus on designing better compression mechanisms; however, adapting the underlying language model to reason effectively over the remaining compressed context remains under-explored. To address this, we propose CAFD (Compressed-Context Adaptation via Full-Context Distillation), a ground-truth-free self-distillation framework that adapts OmniLLMs to fixed compression pipelines without requiring reference answers, rationales, or correctness rewards. CAFD leverages the full-token view of the same multimodal sample as a source of privileged information: a full-context self-teacher provides soft target supervision to a compressed-context student along the student's on-policy trajectory. Evaluated on Qwen2.5-Omni-7B across five audio-video benchmarks, five compression pipelines, and five deployment budgets, CAFD demonstrates consistent gains, improving 120 out of 125 conditions with an average accuracy boost of 1.44 points and recovering 26.9% of the accuracy gap on average. These results demonstrate that the proposed ground-truth-free adaptation offers an effective and practical route to improving the accuracy-efficiency trade-off in deployed OmniLLMs.

发表机构

  • Wuhan University(武汉大学)
  • Didi Chuxing(滴滴出行)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑