基于值强化学习中的批归一化前向模式
On BatchNorm Forward Modes in Value-Based Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
本文证明在离散动作值强化学习中,将批归一化的前向模式从运行统计量切换为批统计量,可显著提升C51和PQN在Atari基准上的性能与稳定性。
中文摘要 AI 辅助
批归一化(BN)显著提升了连续控制中的演员-评论家方法(如CrossQ)的样本效率,然而近期研究报道其在Atari上的离散动作值学习中性能下降。这些失败令人惊讶,因为离散Q网络不存在CrossQ所识别的动作输入分布不匹配问题。我们针对基于目标的C51和无目标的PQN证明,在特定前向传播中简单选择运行统计量或批统计量即可逆转这种性能下降。在C51中,将BN自举前向切换为批统计量模式,其性能显著优于未归一化和LayerNorm基线,并在更新与数据比率高达12时稳定扩展。在PQN中,对动作选择和自举均使用批统计量,可从失败的运行统计量配置中恢复性能。在26个Atari游戏、4亿帧的测试中,该配置的最终总得分高于使用LayerNorm的PQN。我们的结果表明,精心配置的BN可大幅提升离散动作值学习,且其前向协议是算法规范的重要组成部分。
英文摘要
Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies report performance degradation in discrete-action value learning on Atari. These failures are surprising because discrete Q-networks lack the action-input distribution mismatch identified by CrossQ. We show for target-based C51 and target-free PQN that the simple choice between running and batch statistics at specific forward passes can reverse this degradation. In C51, switching the BN bootstrap forward to batch-statistic mode significantly improves performance over unnormalized and LayerNorm baselines and scales stably with update-to-data ratios up to 12. In PQN, using batch-statistics for both action selection and bootstrapping recovers performance from the failing running-statistic configuration. Across 26 Atari games at 400M frames, this configuration achieves a higher final aggregate score than PQN with LayerNorm. Our results show that carefully configured BN can substantially improve discrete-action value learning, and that its forward protocols are an essential part of the algorithm specification.
发表机构
- FAIR at Meta(Meta FAIR)
- Technical University of Darmstadt(达姆施塔特工业大学)
机构由 AI 辅助整理,请以论文原文为准。