arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02819cs.CL

面向全模态推理的以文本为中心的后训练

Text-Centric Post-Training for Omni-Modal Reasoning

Ziyang Cheng, Yuhao Wang, Hongcheng Liu, Qimin Wu, Jingru Fan, Chen Qian, Yanfeng Wang, Yu Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对Omni大模型音视频推理成本高且感知与推理解耦的问题,提出以文本为中心的后训练范式:纯文本训练优化推理,减少数据的音视频RL细化感知,在显著降低算力与数据需求的同时,大幅提升推理并保持感知。

中文摘要 AI 辅助

在Omni大语言模型中改进联合音视频推理通常会产生大量的数据构建和训练成本。我们的诊断揭示了多跳推理的困难,尽管所有对应的单跳问题都能得到正确答案,并表明感知与推理目标的局部优化存在部分解耦。这促使我们在后训练中对这些能力给予不同的侧重。仅文本推理训练在数据来源、模型规模和模型家族上均带来收益。采用性能最佳的纯文本配置,监督微调后接强化学习(RL)使Qwen2.5-Omni-7B的九项推理得分几何平均值较基础模型提升25.83%,超过完整的原生音视频路线,且GPU小时数减少56.6%。完全由纯文本LLM合成的数据训练,在构建和训练中均无音视频数据的情况下,使该几何平均值提升21.01%。然而,纯文本训练会降低感知能力。因此,我们提出一种以文本为中心的后训练范式:纯文本训练提供主要推理优化,而减少数据的原生音视频RL随后细化感知。细化过程使用的输入token比全数据音视频RL少约90%,将感知恢复到基础水平以上,并保留了最佳纯文本流水线推理增益的93.5%。

英文摘要

Improving joint audio-visual reasoning in Omni Large Language Models typically incurs substantial data construction and training costs. Our diagnostics reveal multi-hop reasoning difficulties despite correct answers to all corresponding single-hop questions and suggest partial decoupling in the local optimization of perception and reasoning objectives. This motivates post-training with different emphases on these capabilities. Text-only reasoning training yields gains across data sources, model scales, and families. With the best-performing text-only configuration, supervised fine-tuning followed by reinforcement learning (RL) raises Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% over the base model, outperforming the complete native audio-visual route with 56.6% fewer GPU-hours. Training on data synthesized entirely by a text-only LLM raises this geometric mean by 21.01% without audio-visual data in construction or training. However, text-only training degrades perception. We therefore propose a text-centric post-training paradigm: text-only training provides the main reasoning optimization, and reduced-data native audio-visual RL then refines perception. Refinement uses about 90% fewer input tokens than full-data audio-visual RL, restores perception above the base level, and retains 93.5% of the best-performing text-only pipeline's reasoning gain.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑