发表机构
University of Chicago; University of Florida(芝加哥大学; 佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对并行强化学习中评论家共享导致的价值不匹配问题,提出仅给评论家环境索引的干预方法,在多类任务中提升了学习稳定性与回报,Procgen 游戏中聚合归一化回报提升 40.8%。
AI 中文摘要
当单个策略在同一任务的多个环境(如程序生成的关卡、随机动力学或课程)中并行训练时,现有实现通常会在所有采样环境中使用同一个评论家。然而,不同环境可能会为评论家可见的相同输入分配不同的期望回报。没有环境信息的评论家必须协调不同的价值目标,从而系统性地改变各个环境中的采样优势。我们使用带有多个环境和共同最优臂的说明性老虎机模型,表征这种价值不匹配如何重新分配采样的策略更新,强化无益动作,同时减弱甚至反转有用动作。不使用基线、共享价值或采样环境特定价值的神谕策略,在固定策略下具有相同的平均 logit 更新,并收敛到相同的最优策略,但它们的实际学习路径可能差异很大。该分析提出了一种最小干预:仅向评论家提供记录的环境索引,使其能够分离价值目标。受控的 CartPole 和 MuJoCo 实验揭示了预测的价值偏移、优势和性能差距。在更复杂的 BipedalWalker 和 Procgen 设置中,相同的干预产生了更稳定的学习和更高的回报。在全部 16 个 Procgen 游戏中,多头条件评论家在每个游戏的 600 个未见过关卡上的聚合归一化回报提升了 40.8%。综上,该理论将价值不匹配确定为评论家共享可能降低随机学习动力学的直接机制,该机制无法仅通过标量估计器方差捕获,实验表明,在并行强化学习中,基于索引进行条件化具有广泛有效性。
英文摘要
When a single policy is trained in parallel across multiple environments of the same task, such as procedurally generated levels, randomized dynamics, or curricula, implementations commonly use one critic across all sampled environments. Yet different environments can assign different expected returns to the same input visible to the critic. A critic without environment information must then reconcile distinct value targets, systematically shifting the sampled advantages within individual environments. Using illustrative bandit models with multiple environments and a common optimal arm, we characterize how this value mismatch redistributes sampled policy updates, reinforcing unhelpful actions while attenuating or even reversing useful ones. The oracle processes using no baseline, the shared value, or the value specific to the sampled environment have the same mean logit update at a fixed policy and converge to the same optimal policy, yet their realized learning paths can differ sharply. The analysis motivates a minimal intervention: give only a logged environment index to the critic so that it can separate the value targets. Controlled CartPole and MuJoCo experiments expose the predicted shifted values, advantages, and performance gaps. In the more complex BipedalWalker and Procgen settings, the same intervention yields more stable learning and higher returns. Across all $16$ Procgen games, the multihead conditional critic improves aggregate normalized return on $600$ unseen levels per game by $40.8\%$. In conclusion, the theory identifies value mismatch as a direct mechanism through which critic sharing can degrade stochastic learning dynamics, not captured by scalar estimator variance alone, and the experiments show that conditioning on an index is broadly effective in parallel reinforcement learning.