发表机构
The University of Texas at Austin; UIUC; Emory University; Together AI; Recursive Superintelligence Inc; ELLIS Institute Tübingen(德克萨斯大学奥斯汀分校; 伊利诺伊大学香槟分校; 埃默里大学; 联合人工智能公司; 递归超级智能公司; 图宾根埃利斯研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究 RLVR 中缺失的优化层,提出等谱优化(ISO)框架,包括离线的 ISO-Merger 和在线的 ISO-Optimizer。通过光谱继承,在推理和编码任务中,用更少训练步骤提高准确率,为 RLVR 优化层问题提供了具体解决方案。
AI 中文摘要
可验证奖励的强化学习(RLVR)正在迅速提升语言模型的推理能力,但将奖励反馈转换为权重空间更新的优化层仍未被充分理解。基于先前分析,通过模型权重的奇异结构研究这一缺失层,识别出光谱继承:RLVR 可在获取新行为时重用基础模型的权重光谱。将光谱继承操作化为等谱优化(ISO),这是一个原生 RLVR、固定光谱的优化框架,有离线和在线实例。离线时,ISO-Merger 将共享基础专家的框架变化合并为单个固定光谱模型,无需合并后的数据、展开、梯度更新或策略蒸馏,在无数据合并方法中实现最强聚合性能。在线时,ISO-Optimizer 对框架变量应用选定的基础优化器,保持基础光谱固定。在各种推理和编码任务中,ISO-Optimizer 在报告的运行中提高了准确率,且用更少训练步骤达到匹配分数。总之,ISO 为 RLVR 缺失的优化层提供了具体答案:围绕奖励驱动适应的结构进行训练后设计,继承光谱,优化框架。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through the singular structure of model weights and identify spectral inheritance: RLVR can reuse the base model's weight spectra while acquiring new behavior through changes in the associated input and output singular frames. We operationalize spectral inheritance as Isospectral Optimization (ISO), an RLVR-native, fixed-spectrum optimization framework with complementary offline and online instantiations. Offline, ISO-Merger combines the frame changes of shared-base specialists into a single fixed-spectrum model, requiring no post-merge data, rollouts, gradient updates, or on-policy distillation (OPD). It recovers complementary specialist capabilities and achieves the strongest aggregate performance among the compared data-free merging methods. Online, ISO-Optimizer applies a chosen base optimizer, including AdamW and Muon, to the frame variables while keeping the base spectra fixed. Across reasoning and coding tasks ranging from 1.5B to 8B parameters, ISO-Optimizer improves accuracy in the reported runs and reaches matched scores with substantially fewer training steps. On Qwen3-8B-Base, AdamW reaches an aggregate accuracy of 0.495 after 270 training steps. ISO-AdamW reaches the same accuracy after only 100 training steps and improves further to 0.509 after 210 training steps. Together, ISO offers a concrete answer to RLVR's missing optimization layer: rather than inheriting pre-training optimization wholesale, design post-training around the structure of reward-driven adaptation: inherit the spectrum, optimize the frames.
CommentsPreprint