arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用你所说的语言推理:为VideoLLM中的推理蒸馏进行个人习语式的离线策略轨迹改写

Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs

Ji Soo Lee, Jinyoung Park, Seohyun Lee, Jongha Kim, Joonmyung Choi, Jinsung Yoon, Hyunwoo J. Kim

arXiv 2608.26684首次发表:更新:

发表机构

KAIST; Korea University; Google Cloud AI Research(韩国科学技术院; 高丽大学; 谷歌云人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对VideoLLM的推理蒸馏提出Echo-GRPO框架,通过改写教师特权轨迹为学生自身习语实现策略对齐,实例化的VideoEcho-R1在多骨干和基准测试中均获性能提升,且作为插件模块适配多种训练框架。

AI 中文摘要

近期大型语言模型在复杂推理任务上表现出色,其中采用组相对策略优化(GRPO)的强化学习已成为在自生成轨迹上优化模型的主流范式。然而,GRPO的在线策略属性限制了模型只能利用自身已具备的推理技能,无法学习更高级的能力。现有研究会注入来自更强教师策略的特权推理轨迹来指导训练,但这些轨迹与学生策略本质上属于分布外数据。我们发现,在线策略与离线策略之间的这种不匹配会导致语义关键的推理标记出现梯度裁剪,最终仅奖励正确答案,却未学习到支撑答案的推理过程。因此,我们提出了Echo-GRPO框架,该框架让模型用自身生成的语言进行推理。Echo-GRPO并非模仿教师模型的低概率特权轨迹,而是通过双参考解码在保留语义的同时,将这些轨迹改写为学生策略自身的个人习语,即其特有的词汇和表达模式。我们将该框架实例化为用于视频推理蒸馏的VideoEcho-R1,在三个多模态LLM骨干和五个基准测试中均实现了持续的性能提升。最后,我们证明个人习语式改写是一个插件模块,可持续改进用于推理蒸馏的强化学习和监督微调框架,表明策略对齐的监督不仅适用于GRPO。

英文摘要

Recent large language models achieve strong performance on complex reasoning tasks, where reinforcement learning with Group Relative Policy Optimization (GRPO) has emerged as a leading paradigm for optimizing models on self-generated trajectories. However, the on-policy nature of GRPO bounds the model to the reasoning skills it can already produce, restricting to learn more advanced capabilities. Prior works inject privileged reasoning traces from a stronger teacher policy to guide training, yet these traces are inherently out of distribution with respect to the student policy. We observe that this mismatch between on-policy and off-policy causes gradient clipping on semantically critical reasoning tokens, ultimately rewarding correct answers while leaving the reasoning that justifies them unlearned. Hence, we propose \textbf{Echo-GRPO}, a framework that lets the model reason in the words it speaks. Rather than imitating low-probability privileged traces from the teacher model, Echo-GRPO rewrites them into the student policy's own \textit{idiolect}, that is, its own characteristic vocabulary and expression patterns, while preserving their semantics via Dual-Reference Decoding. We instantiate this framework as \textbf{VideoEcho-R1} for video reasoning distillation, achieving consistent improvements across three multimodal LLM backbones and five benchmarks. Finally, we show that our idiolectal paraphrasing is a plug-in module that consistently improves both RL and supervised fine-tuning frameworks for reasoning distillation, demonstrating that policy-aligned supervision extends beyond GRPO.

CommentsAccepted to NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑