arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一步一引导:通过跨步控制缓解多域强化学习中的高阶干扰

One Step, One Lead: Mitigating Higher-Order Interference in Multi-Domain Reinforcement Learning via Cross-Step Control

Zihan Lin, Xiaohan Wang, Jie Cao, Jiajun Chai, Guojun Yin, Wei Lin, Ran He

arXiv 2609.06469首次发表:更新:

发表机构

MAIS&NLPR, Institute of Automation, Chinese Academy of Sciences; Meituan; School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统全国重点实验室&模式识别国家重点实验室; 美团; 中国科学院大学前沿交叉科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出OSOL方法,通过跨步控制缓解多域强化学习中的高阶干扰,利用相邻检查点的令牌足迹排序回弹风险并自适应修正,在Qwen3-30B-A3B上提升5.7%。

AI 中文摘要

跨多个领域的强化学习(RL)可以拓宽大型语言模型(LLM)的推理能力,但联合训练常常会降低单个领域的性能,并可能使优化过程不稳定。现有工作通常使用一阶梯度对齐或基于曲率的代理,从单步视角来诊断这种干扰。我们表明,这种视角可能会遗漏一种关键的序列干扰形式:即使连续的已实现更新在输出空间中部分地相互逆转,同一点上的域梯度仍可能保持近似正交。我们进一步表明,连续的令牌对数概率足迹可以直接从相邻检查点恢复这种交互,作为输出空间中的局部二阶交互,而无需显式重建同一步的曲率。基于这一见解,我们提出了OSOL,它在每次迭代中指定一个焦点域,使用前一个检查点的足迹来对令牌级别的回弹风险进行排序,并在标准GRPO更新中应用一种按漂移排序、自适应缩放的修正。我们的分析表明,这种修正抑制了目标跨步输出回溯分量。对照研究进一步表明,跨步回溯与后续任务损害的相关性比同一点梯度诊断更强,而前一个足迹对未来回弹风险的排序比基于Hessian的代理更准确。在Qwen3-30B-A3B上,OSOL达到了0.4822的域宏平均分数,比最强的对比基线提高了5.7%,且无需显式的高阶微分。

英文摘要

Reinforcement learning (RL) across multiple domains can broaden the reasoning capabilities of large language models (LLMs), yet joint training often degrades individual-domain performance and can destabilize optimization. Existing work typically diagnoses such interference from a single-step view using first-order gradient alignment or curvature-based proxies. We show that this view can miss a critical form of sequential interference: same-point domain gradients may remain nearly orthogonal even when consecutive realized updates partially reverse one another in output space. We further show that consecutive token log-probability footprints recover this interaction directly from adjacent checkpoints as a local second-order interaction in output space, without explicitly reconstructing same-step curvature. Building on this insight, we propose OSOL, which designates a focus domain at each iteration, uses the preceding checkpoint footprint to rank token-level rebound risk, and applies a drift-ranked, adaptively scaled correction within the standard GRPO update. Our analysis shows that this correction suppresses the targeted cross-step output backtracking component. Controlled studies further show that cross-step backtracking is more strongly associated with subsequent task damage than same-point gradient diagnostics, while the preceding footprint ranks future rebound risk more accurately than Hessian-based proxies. On Qwen3-30B-A3B, OSOL reaches a domain-macro average of 0.4822, improving by 5.7% over the strongest compared baseline, without explicit higher-order differentiation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑