arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24479cs.LG

WarpSAC:通过重新思考探索与利用,迈向可拓展离线强化学习的顶峰

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

Zihao Wu, Hongyao Tang, Yi Ma, Huizhong Song, Pengyi Li, Yifu Yuan, Fei Ni, Jinyi Liu, Wei Wei, Jianrong Wang, Yan Zheng, Jianye Hao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对大规模并行模拟下离线RL的数据 regime 变化问题,提出具 regime 感知的WarpSAC算法家族,适配不同数据规模场景,在多基准任务及现实部署上较FlashSAC实现显著性能提升。

中文摘要 AI 辅助

大规模并行模拟改变了离线强化学习(RL)的训练数据 regime,对为数据受限回放设计的稳定器构成挑战。通过在8个基准系列上开展受控实验,我们发现这些稳定器具有数据 regime 依赖性:参数归一化有助于回放覆盖范围较窄的场景,但在数据充足时会限制价值拟合;而在高吞吐量操作场景中,可放宽截断双Q值的使用。年龄偏置回放加权可提升各 regime 下的学习效率,尤其在网络容量有限时效果显著。基于这些发现,我们提出了WarpSAC,这是一个具备 regime 感知能力的离线RL算法家族。WarpSAC采用样本权重衰减实现高效利用,并提供两个变体:WarpSAC-L(归一化开启、截断双Q值)适用于数据受限的CPU规模训练,WarpSAC-A(归一化关闭、单Q值)适用于数据充足的GPU并行训练。WarpSAC在9个CPU规模环境上的归一化分数-步AUC较FlashSAC提升4.5%,在14个GPU并行环境上提升23.1%;将UnitreeG1TransportBox-v1的成功率从19.8%提升至96.4%,使MuJoCo Playground上的平均归一化挂钟时间AUC提升19.1%,在Unitree G1上的现实部署速度较FlashSAC快36.4%。这些结果表明,可拓展离线RL应根据可用数据 regime 调整其稳定器。

英文摘要

Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for data-limited replay. Through controlled experiments across eight benchmark families, we show that these stabilizers are data-regime-dependent: parameter normalization helps with narrow replay coverage but restricts value fitting when data are abundant, while clipped double-Q can be relaxed in high-throughput manipulation. Age-biased replay weighting improves learning efficiency across regimes, especially with limited network capacity. Based on these findings, we propose WarpSAC, a regime-aware family of off-policy RL algorithms. WarpSAC uses Sample Weight Decay for efficient exploitation and provides two variants: WarpSAC-L (Norm ON, clipped double-Q) for data-limited CPU-scale training, and WarpSAC-A (Norm OFF, single-Q) for data-abundant GPU-parallel training. WarpSAC improves normalized score--step AUC over FlashSAC by 4.5% across nine CPU-scale environments and 23.1% across fourteen GPU-parallel environments. It increases UnitreeG1TransportBox-v1 success rate from 19.8% to 96.4%, improves mean normalized wall-time AUC on MuJoCo Playground by 19.1%, and achieves 36.4% faster sim-to-real deployment on Unitree G1 than FlashSAC. These results show that scalable off-policy RL should adapt its stabilizers to the available data regime.

发表机构

  • Tianjin University(天津大学)
  • Shanxi University(山西大学)
  • Imperial College London(伦敦帝国学院)

机构由 AI 辅助整理,请以论文原文为准。

↑