arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

控制多样化强化微调:解耦RL后训练的共享控制瓶颈

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training

Binwen Tan, Jingchao Wang, Dengzhe Hou, Lingyu Jiang, Zeyuan Wu, Yunhan Shen, Fangzhou Lin, Kazunori Yamada, Atsushi Koike

arXiv 2608.08224首次发表:更新:

AI 中文总结

本研究提出CD-RFT方法,定义后训练控制系数揭示RL后训练的共享控制瓶颈,通过正则化减少控制坍缩,在Qwen2.5-7B等模型上实现控制解耦与多任务能力提升,效果可迁移至Llama-3.2-3B。

AI 中文摘要

强化学习后训练(RL post-training)可解锁大语言模型(LLM)的复杂推理能力,但基准分数仅能体现模型是否提升,无法反映其内部变化,也无法体现有限能力在不同任务间的分配方式。一条具代表性的可解释性研究路线将RL微调的成功归因于更强、更多样的回路激活。本研究通过将激活与控制分离,对这一以激活为核心的解释提出挑战:激活的回路未必能控制后训练的奖励增益。本研究借鉴代谢控制分析(Metabolic Control Analysis),定义后训练控制系数(Post-training Control Coefficient)以衡量各组件对奖励增益的控制作用,并按任务家族将这些系数排列为控制矩阵,同时搭配激活幅度矩阵。本研究将跨任务控制的集中现象称为共享控制瓶颈(Shared Control Bottleneck),将激活集中与控制集中的差异称为激活-控制差距(Activation-Control Gap)。这一概念揭示:高度共享的激活可与任务特异性控制共存,而较小的差距则表明控制已坍缩至共享方向、丧失任务特异性。为减少这种坍缩,本研究在训练损失中对共享控制瓶颈进行正则化,并提出控制多样化强化微调(Control-Diverse Reinforcement Fine-Tuning, CD-RFT)。精确的正则化梯度需二阶自动微分,与Flash Attention不兼容,因此本研究推导了一阶代理,其最坏情况开销低于8%。在Qwen2.5-7B模型上,CD-RFT实现了最大程度的控制解耦,且在数学、代码、逻辑任务上的多任务能力优于匹配的GRPO方法;无KL惩罚的变体在pass@1指标上领先,带KL惩罚的变体则在大k的pass@k覆盖率上领先,而KL原本会降低该指标。综合来看,这些结果表明,共享控制瓶颈既是机制层面的诊断工具,也是训练正则化器,且控制解耦与能力提升可迁移至Llama-3.2-3B模型。

英文摘要

Reinforcement learning post-training unlocks complex reasoning in LLMs. Yet benchmark scores reveal only whether a model improved, not what changed inside it, nor how it splits finite capability across tasks. A representative interpretability line attributes the success of RL fine-tuning to stronger and more diverse circuit activation. We challenge this activation-centered account by separating activation from control: an activated circuit need not control the post-training reward gain. Adapting Metabolic Control Analysis, we define the Post-training Control Coefficient to measure component control over reward gain and arrange these coefficients by task family into a control matrix, paired with an activation-magnitude matrix. We call cross-task control concentration the Shared Control Bottleneck and the difference between activation and control concentration the Activation-Control Gap. This reveals that highly shared activations can coexist with task-specific control, while a small gap indicates that control has collapsed onto a shared direction and lost task specificity. To reduce this collapse, we regularize the post-training loss with the Shared Control Bottleneck and propose Control-Diverse Reinforcement Fine-Tuning (CD-RFT). The exact regularizer gradient requires second-order automatic differentiation incompatible with flash attention, so we derive a first-order proxy with worst-case overhead below eight percent. On Qwen2.5-7B, CD-RFT achieves the largest control decoupling and improves multi-task capability over matched GRPO across mathematics, code, and logic. The no-KL variant leads on pass@1, and the KL-penalized variant leads on large-k pass@k coverage that KL otherwise degrades. Together, these results show that the Shared Control Bottleneck is both a mechanistic diagnostic and a training regularizer, and that control decoupling and capability gains transfer to Llama-3.2-3B.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑