arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于RLVR后训练中的语言漂移

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

arXiv 2610.02015首次发表:更新:

发表机构

Saarland University(萨尔兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文证明RLVR后训练中语言漂移由优化压力导致且无法在不牺牲性能的前提下约束,并实验验证其仅在新颖推理任务中出现。

AI 中文摘要

近期,基于可验证奖励的强化学习后训练(RLVR)范式,大语言模型推理模型取得了显著进展,使其能够完成极其复杂的任务。然而,随着能力提升,大语言模型在其思维链(CoTs)中日益表现出语言漂移的迹象:即使用不寻常、非标准且看似无意义的语言。尽管这一现象已被充分记录,并可能损害思维链的可监控性,但其成因至今仍知之甚少。本文中,我们识别了语言漂移发生的条件:我们从理论上证明了RLVR的优化压力允许无界语言漂移,而监督微调则不会。随后,我们通过实验表明,语言漂移尤其会在RLVR处理新型推理任务时出现——即当目标行为无法从基础模型中引出时。最后,我们证明了在不约束预期奖励的情况下,无法约束语言漂移,这表明在前沿RLVR后训练期间,若不损害性能,思维链的可监控性就无法提升。

英文摘要

Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift in their chains of thought (CoTs): unusual, non-standard, and seemingly nonsensical language use. Although it is well-documented---and can potentially impair CoT monitorability---the causes of language drift are thus far poorly understood. In this paper, we identify the conditions under which language drift occurs: we prove theoretically that RLVR optimization pressure permits unbounded language drift, while supervised fine-tuning does not. We then show empirically that language drift specifically arises during RLVR on novel reasoning tasks---i.e. when the target behavior cannot be drawn out of the base model. Finally, we prove that it is not possible to constrain language drift without constraining expected reward, suggesting that CoT monitorability cannot be improved without harming performance during RLVR post-training at the frontier.

Comments22 pages; 15 figures; 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑