arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

策略即数据:基于优势回归的重放策略对偶平均

Policy as Data: Replay-Based Policy Dual Averaging via Advantage Regression

Nianli Peng, Geoffrey J. Gordon, Kianté Brantley

arXiv 2610.04638首次发表:更新:

发表机构

Harvard University; Carnegie Mellon University(哈佛大学; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出RDA2C算法,将重放数据作为对偶目标,通过优势回归累积策略改进,在MuJoCo和Atari基准上超越PPO等基线。

AI 中文摘要

演员-评论家方法通过重用过去的经验来提高样本效率。然而,历史数据通常被视为当前策略改进更新的离策略样本。这项工作引入了正则化对偶平均演员-评论家(RDA2C),为重放赋予了不同的角色。在正则化对偶平均中,后续策略由累积的策略改进反馈决定,因此历史优势估计直接贡献于演员目标,而不仅仅是最新更新。RDA2C存储带有评论家估计优势标签的状态-动作样本,将聚合数据集拟合到对偶得分模型$Z_\ heta$,并使用熵镜像映射从累积得分模型推导当前策略。这样,重放定义了一个经验对偶目标,策略从中计算得出。为了分析RDA2C,我们建立了一个有限时间价值差距分解,将正则化对偶平均项与由于陈旧重放监督拟合、评论家偏差、有限缓冲区方差和重放覆盖导致的误差分开,并陈述了每个误差有界的假设。RDA2C接受来自任何评论家的优势标签。使用GAE标签,RDA2C在八个MuJoCo任务中的六个和十二个Atari游戏中的八个上优于PPO。RDA2C还在八个MuJoCo任务中的六个上优于最接近的对偶平均基线AAPDA。使用孪生$Q$标签,在匹配的批大小和更新频率下,RDA2C与SAC相当。

英文摘要

Actor-critic methods reuse past experience to improve sample efficiency. However, historical data are typically regarded as off-policy samples for the current policy-improvement update. This work introduces Regularized Dual Averaging Actor Critic (RDA2C), which assigns a distinct role to replay. In regularized dual averaging, the subsequent policy is determined by accumulated policy-improvement feedback, so historical advantage estimates contribute directly to the actor objective rather than solely to the most recent update. RDA2C stores state-action samples with critic-estimated advantage labels, fits a dual score model $Z_θ$ to the aggregated dataset, and derives the current policy from the accumulated score model using the entropy mirror map. In this way, replay defines an empirical dual objective from which the policy is computed. To analyze RDA2C, we establish a finite-time value-gap decomposition, separating the regularized dual-averaging term from errors due to stale-replay supervised fitting, critic bias, finite-buffer variance, and replay coverage, and stating the assumptions under which each error is bounded. RDA2C accepts advantage labels from any critic. With GAE labels, RDA2C outperforms PPO on six of eight MuJoCo tasks and eight of twelve Atari games. RDA2C also outperforms AAPDA, the closest dual-averaging baseline, on six of eight MuJoCo tasks. With twin-$Q$ labels, RDA2C matches SAC at matched batch size and update frequency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑