arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10481cs.LGcs.AIcs.CL

ARMOR:使用离策略锚样本稳定在线大语言模型强化学习

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples

Kexin Huang, Junkang Wu, Jinda Lu, Shuo Yang, Chiyu Ma, Jiancan Wu, Xiang Wang, Xiangnan He, Guoyin Wang, Jingren Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型强化学习训练不稳定问题,提出ARMOR框架,通过锚展开利用离策略数据保留解决方案模式,混合优化重新制定策略目标实现可控探索,经实验验证可有效减轻验证崩溃,提升性能。

中文摘要 AI 辅助

强化学习(RL)显著提升了大语言模型(LLMs)的推理能力,但训练过程仍然非常脆弱。本文研究了这种不稳定性的一个关键来源:过度优化,即模型利用训练启发式方法而牺牲了可泛化推理能力。虽然反向KL正则化是防止这种退化的标准防御方法,但分析表明它在此情况下往往不足,因为它无法确保参考分布的全面覆盖。为解决此问题,我们提出了ARMOR(用于RL的锚展开和混合优化)框架,该框架将范式从被动惩罚转变为主动样本稳定。ARMOR包含两个关键组件:(1)锚展开,利用来自参考策略的离策略数据来保留已建立的解决方案模式;(2)混合优化,重新制定策略目标以实现可控探索而不依赖辅助损失。在推理基准上的大量实验验证了ARMOR有效地减轻了验证崩溃,在延长的训练范围内实现了持续的性能提升。

英文摘要

Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), yet the training process remains notoriously fragile. In this work, we investigate a critical source of this instability: over-optimization, where models exploit training heuristics at the expense of generalizable reasoning. While reverse KL regularization is the standard defense against such degradation, our analysis reveals that it is often insufficient in this regime, as it fails to ensure comprehensive coverage of the reference distribution. To address this, we propose ARMOR (Anchor Rollout and Mixed Optimization for RL), a framework that shifts the paradigm from passive penalty to active sample stabilization. ARMOR comprises two key components: (1) Anchor Rollout, which leverages off-policy data from the reference policy to preserve established solution patterns; and (2) Mixed Optimization, which reformulates the policy objective to enable controlled exploration without relying on auxiliary losses. Extensive experiments on reasoning benchmarks validate that ARMOR effectively mitigates validation collapse, enabling sustained performance improvements over extended training horizons.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Peking University(北京大学)
  • Dartmouth College(达特茅斯学院)

机构由 AI 辅助整理,请以论文原文为准。

↑