arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24144cs.LG

运气非技能:配对轨迹何时有助于LLM智能体的组相对强化学习?

Luck Is Not Skill: When Do Paired Rollouts Help Group-Relative RL of LLM Agents?

  • BRAC University(BRAC大学)

机构由 AI 辅助整理,请以论文原文为准。

Nazmus Sakib

AI总结:

本研究探讨配对轨迹在组相对强化学习中的作用,发现其可降低奖励对比方差但未必降低梯度方差,实验显示配对在工具故障下提升成功率,但未显著改善学习速度。

AI中文摘要:

组相对强化学习比较同一提示下的轨迹,但独立的环境噪声可能干扰这些比较。我们研究了配对轨迹,它在每组内共享一个事件键控的噪声调度,同时保留每条轨迹的边缘分布。配对消除了奖励对比方差中调度间成分,但未必能降低梯度方差。对于单侧评分器噪声,我们推导了降低的精确条件,并给出了一个反例,其中奖励对比改善而梯度方差增加。一项受控研究在工具故障和评分器翻转条件下训练了一个2B工具使用智能体,每种设计使用三个随机种子。该协议已注册,并包含一个已公开的先前完成的试点。在工具故障下,配对将最终噪声测试成功率平均提高了+5.1个百分点,三个种子的差异均为正,但未达到注册的学习曲线标准。在评分器翻转下也未达到该标准:验证AUC差异为+0.003(95%区间[-0.029, +0.033])。对来自两个故障训练轨迹的八个不同检查点进行的梯度探针发现,在两种噪声类型下,平均中心化协方差迹均较低:评分器翻转降低21%至30%,工具故障降低40%至63%。这些有限样本测量支持方差机制,但未确立普遍的学习速度优势。结果区分了改善奖励比较、降低估计器方差和改善学习。

英文摘要:

Group-relative reinforcement learning compares rollouts of the same prompt, but independent environment noise can obscure these comparisons. We study paired rollouts, which share an event-keyed noise schedule within each group while preserving each rollout's marginal distribution. Pairing removes the between-schedule component of reward-contrast variance, but need not reduce gradient variance. For one-sided grader noise, we derive an exact condition for reduction and give a counterexample in which reward contrasts improve while gradient variance increases. A controlled study trains a 2B tool-use agent under tool faults and grader flips, with three seeds per design. The protocol was registered with a disclosed, previously completed pilot. Under tool faults, pairing improves final noisy-test success by +5.1 percentage points on average, with all three seed differences positive, but misses the registered learning-curve criterion. The criterion is also missed under grader flips: the validation-AUC difference is +0.003 (95% interval [-0.029, +0.033]). A gradient probe on eight distinct checkpoints from two fault-trained trajectories finds lower mean-centered covariance traces under both noise types: 21 to 30% for grader flips and 40 to 63% for tool faults. These finite-sample measurements support the variance mechanism without establishing a general learning-speed benefit. The results distinguish improving reward comparisons, reducing estimator variance, and improving learning.

补充信息

↑