arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34344cs.LGcs.AI

学习转向,转向以见:通过可训练向量揭示大型语言模型中RLVR的几何结构

Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors

Yuchen Cai, Ding Cao, Qixiang Yin, Xin Xu, Kai Yang, Siye Wu, Pengyuan Wang, Jiaxuan Wang, Weijie Liu, Saiyong Yang, Guangzhong Sun, Guiquan Liu, Junfeng Fang

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过向量转向揭示RLVR在LLM中的低维几何特性,提出Alpha-Stabler框架稳定训练并提升RL增益,实验验证于5个LLM和6个任务。

中文摘要 AI 辅助

强化学习(RL)已成为增强大型语言模型推理能力的关键范式,然而参数更新的高维性使其训练动态难以分析。我们研究了具有可验证奖励的强化学习(RLVR),并利用向量转向来识别激活空间中与RL诱导增益相关的低维有效流形。我们揭示了两个几何性质。(1)有效流形容量:重现RL增益所需的容量可以非常小,但并非无限可压缩;在极低容量下,干预维度和输入依赖性表达能力成为关键约束,且该需求随注入深度变化。(2)控制流形分离:有效控制方向主要位于激活主子空间的低方差补集中。在任务和基础模型内,学习到的几何结构在训练配置间基本保持一致,跨任务时几何对齐与能力迁移相关。在5个LLM和6个可验证奖励任务上的实验支持了这些发现。随后,我们提出了Alpha-Stabler,一个即插即用框架,包含一个预测器(Predictor)用于监测主子空间侵入以进行早期崩溃预警,以及一个控制器(Controller)在后向传播期间移除激活梯度的主子空间分量,同时保留正交补集。Alpha-Stabler稳定了2000步的训练,并持续提升RL增益,为稳健的后训练提供了实用见解。代码:此https URL

英文摘要

Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze. We study reinforcement learning with verifiable rewards (RLVR) and use vector steering to identify a low-dimensional effective manifold in activation space associated with RL-induced gains. We uncover two geometric properties. (1) Effective Manifold Capacity: the capacity needed to reproduce RL gains can be very small but is not infinitely compressible; at extremely low capacity, intervention dimensionality and input-dependent expressiveness become key constraints, and this requirement varies with injection depth. (2) Control Manifold Separation: effective control directions lie mainly in the low-variance complement of the activation principal subspace. Within a task and base model, the learned geometry stays largely consistent across training configurations, and across tasks geometric alignment correlates with capability transfer. Experiments on 5 LLMs and 6 verifiable-reward tasks support these findings. We then propose Alpha-Stabler, a plug-and-play framework with a Predictor that monitors principal-subspace intrusion for early collapse warnings, and a Controller that removes the principal-subspace component of activation gradients during backpropagation while preserving the orthogonal complement. Alpha-Stabler stabilizes training for 2,000 steps and consistently improves RL gains, offering practical insights for robust post-training. Code: https://github.com/caiyuchen-ustc/On_Policy_Vector_Training

发表机构

  • USTC(中国科学技术大学)
  • Tencent Hunyuan(腾讯混元)
  • BUPT(北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑