arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21273cs.LG

奖励通道中的暗室:密集预测奖励使GRPO训练的大语言模型智能体崩溃——以及实际有效的方法

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works

Yu Wang

首次发表
浏览论文内容

中文总结 AI 辅助

研究发现密集预测奖励在GRPO训练下会使大语言模型智能体崩溃,通过单因素消融定位原因,提出方差分布标准,还通过受控信号传递矩阵对比奖励通道和辅助损失通道,表明奖励通道至多中性,辅助损失通道有优势。

中文摘要 AI 辅助

密集的逐步骤监督是稀疏奖励、长视野大语言模型智能体的一种有吸引力的补救方法:为智能体预测其下一个观察结果给予奖励,记忆也应随之而来。我们表明,在组归一化强化学习(GRPO)下,这种方法不仅失败——它还会破坏策略。在ALFWorld上的Qwen3 - 1.7B/4B/8B中,基于势能的预测奖励会使每次运行陷入退化吸收状态(预测准确率 -> 1.0,任务成功率 -> 0,情节长度固定在视野):即“暗室”病态,由优化器自动构建。单因素消融定位了原因——仅去除GRPO的标准差归一化就能使相同奖励从灾难性(零成功率)变为基线水平——一个两行命题解释了原因:在全失败组中,z分数优势对塑造系数不变,因此有界奖励变成无界压力且退火无济于事。我们的核心见解对此进行了推广:z分数放大的是密集信号在组内的方差,而全失败组占主导,所以方差因掌握而衰减的信号在结构上是不同的。这种方差分布标准追溯了我们的崩溃情况,对尚未运行的分支进行了预注册预测,并与已发表的奖励通道成功案例一致(是兼容性检查,而非独立测试)。最后,一个受控信号传递矩阵(相同信号,仅改变消耗机制)表明奖励通道至多是中性的,而辅助损失通道获得约20分——并且随机排列的金币安慰剂与真金币分支匹配,所以即使没有正确标签,差距依然存在。端点是单种子的;种子复制和组大小控制已预注册且正在进行中。

英文摘要

Dense per-step supervision is the standard remedy for sparse-reward long-horizon LLM agents: reward the policy for predicting its next observation, which looks provably safe under potential-based shaping. Published prediction-reward and auxiliary-loss variants report both successes and instabilities; we supply the controlled account: 74 preregistered arms dissect one fixed prediction signal under GRPO across ALFWorld, WebShop, a synthetic POMDP, and Qwen3-1.7B/4B/8B, varying only the delivery mechanism. (1) Every run sustaining this difference-form reward under untouched std normalization (no filtering, dynamic-sampling, or decoupling mitigations) collapses: eleven runs across scales, coefficients, group sizes, and groupings (the floor-bound synthetic environment stalls instead); ALFWorld runs end in an absorbing state (prediction accuracy -> 1.0, success -> 0): the optimizer builds the "dark room". The algebra is one line: in all-fail groups z-scoring cancels the shaping coefficient; removing only std normalization restores baseline parity. (2) A signal's danger is set by its within-group variance trajectory, plus hackability as a second axis; it retrodicts every reward-channel collapse and survives preregistered prospective tests. (3) The same signal as a teacher-forced auxiliary loss is harmless on ALFWorld at 4B, but the gain is not the signal's: content-free placebos as a class match or beat gold at both matched seeds (s0: 78.8 vs 68.6; s42: 67.9 vs 57.9); the auxiliary update is the regularizer. (4) At 8B the recipe turns bistable: gold full-weight locks two of three seeds; every content-free or reduced-weight arm stays healthy. No ALFWorld or WebShop reward-channel variant measurably beats its matched-normalization baseline and no gold signal measurably outperforms its content-free placebo: the delivery channel, not the content, decides; which channel is safe is regime-dependent.

补充信息

↑