arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReBRAC-v2:王者归来

ReBRAC-v2: The Return of the King

Denis Tarasov, Robert K. Katzschmann

arXiv 2608.01205首次发表:更新:

AI 中文总结

本文提出ReBRAC-v2,通过改进传统行为正则化演员-评论家算法,在OGBench、D4RL等基准测试中取得最优离线强化学习性能,证明严谨可迁移的工程设计可实现出色表现。

AI 中文摘要

近期离线强化学习方法愈发依赖表达能力强的生成策略与专门的价值引导机制。本文探究能否通过系统地改进传统的行为正则化演员-评论家算法,同时保留其算法简洁性,取得相当的进展。我们提出ReBRAC-v2,该算法直接训练精确似然归一化流作为强化学习的演员,结合似然、均方误差(MSE)和平均绝对误差(MAE)行为正则化,并集成基于分类的残差评论家、分阶段优化以及多样本测试时动作选择。我们未针对每个任务单独调整该方案,而是在6个具有挑战性的OGBench任务上通过约600次贝叶斯建议开发出单一共享配置,冻结所有结构与优化选择,仅在16点网格上调整2个行为正则化系数。在10个常见的基于状态的OGBench类别中,ReBRAC-v2的平均得分为74.8,次优聚合结果为52.3,且在8个类别中排名第一。相同方案无需结构改动,在D4RL AntMaze(90.2)和Adroit(33.6)的对比中取得最强平均结果。固定方案的消融实验显示,该算法对混合克隆目标、分阶段训练、充足的流容量以及多样本推理最为敏感,同时表明若干较小的选择依赖于其他超参数的值。这些结果表明,严谨且可迁移的工程设计无需摒弃极简的离线强化学习基础,即可实现最先进的聚合性能。

英文摘要

Recent offline reinforcement learning methods increasingly rely on expressive generative policies and specialized value-guidance mechanisms. We ask whether comparable progress can instead come from systematically modernizing a conventional behavior-regularized actor-critic while preserving its algorithmic simplicity. We introduce ReBRAC-v2, which directly trains an exact-likelihood normalizing flow as the RL actor, combines likelihood, MSE, and MAE behavior regularization, and integrates a classification-based residual critic, staged optimization, and multi-sample test-time action selection. Rather than tuning this recipe separately for every task, we develop a single shared configuration via roughly 600 Bayesian proposals on six challenging OGBench tasks, freeze all structural and optimization choices, and adapt only two behavior-regularization coefficients over a 16-point grid. Across ten common state-based OGBench categories, ReBRAC-v2 averages 74.8 compared to 52.3 for the next-best aggregate result and ranks first in eight categories. The same recipe, without structural changes, obtains the strongest averages in our comparisons on D4RL AntMaze (90.2) and Adroit (33.6). Fixed-recipe ablations show the largest sensitivity to the selected mixed cloning objective, staged training, sufficient flow capacity, and multi-sample inference, while showing that several smaller choices depend on the values of other hyperparameters. These results show that disciplined, transferable engineering can achieve state-of-the-art aggregate performance without abandoning a minimalist offline RL foundation.

Commentshttps://github.com/DT6A/ReBRAC-v2

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑