WarpSAC:通过重新思考探索与利用,迈向可拓展离线强化学习的顶峰
WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation
浏览论文内容
中文总结 AI 辅助
该研究针对大规模并行模拟下离线RL的数据 regime 变化问题,提出具 regime 感知的WarpSAC算法家族,适配不同数据规模场景,在多基准任务及现实部署上较FlashSAC实现显著性能提升。
中文摘要 AI 辅助
大规模并行模拟改变了离线强化学习(RL)的训练数据 regime,对为数据受限回放设计的稳定器构成挑战。通过在8个基准系列上开展受控实验,我们发现这些稳定器具有数据 regime 依赖性:参数归一化有助于回放覆盖范围较窄的场景,但在数据充足时会限制价值拟合;而在高吞吐量操作场景中,可放宽截断双Q值的使用。年龄偏置回放加权可提升各 regime 下的学习效率,尤其在网络容量有限时效果显著。基于这些发现,我们提出了WarpSAC,这是一个具备 regime 感知能力的离线RL算法家族。WarpSAC采用样本权重衰减实现高效利用,并提供两个变体:WarpSAC-L(归一化开启、截断双Q值)适用于数据受限的CPU规模训练,WarpSAC-A(归一化关闭、单Q值)适用于数据充足的GPU并行训练。WarpSAC在9个CPU规模环境上的归一化分数-步AUC较FlashSAC提升4.5%,在14个GPU并行环境上提升23.1%;将UnitreeG1TransportBox-v1的成功率从19.8%提升至96.4%,使MuJoCo Playground上的平均归一化挂钟时间AUC提升19.1%,在Unitree G1上的现实部署速度较FlashSAC快36.4%。这些结果表明,可拓展离线RL应根据可用数据 regime 调整其稳定器。
英文摘要
Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for data-limited replay. Through controlled experiments across eight benchmark families, we show that these stabilizers are data-regime-dependent: parameter normalization helps with narrow replay coverage but restricts value fitting when data are abundant, while clipped double-Q can be relaxed in high-throughput manipulation. Age-biased replay weighting improves learning efficiency across regimes, especially with limited network capacity. Based on these findings, we propose WarpSAC, a regime-aware family of off-policy RL algorithms. WarpSAC uses Sample Weight Decay for efficient exploitation and provides two variants: WarpSAC-L (Norm ON, clipped double-Q) for data-limited CPU-scale training, and WarpSAC-A (Norm OFF, single-Q) for data-abundant GPU-parallel training. WarpSAC improves normalized score--step AUC over FlashSAC by 4.5% across nine CPU-scale environments and 23.1% across fourteen GPU-parallel environments. It increases UnitreeG1TransportBox-v1 success rate from 19.8% to 96.4%, improves mean normalized wall-time AUC on MuJoCo Playground by 19.1%, and achieves 36.4% faster sim-to-real deployment on Unitree G1 than FlashSAC. These results show that scalable off-policy RL should adapt its stabilizers to the available data regime.
发表机构
- Tianjin University(天津大学)
- Shanxi University(山西大学)
- Imperial College London(伦敦帝国学院)
机构由 AI 辅助整理,请以论文原文为准。