arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多保真策略梯度稳定数据稀缺强化学习

Multi-Fidelity Policy Gradients Stabilize Data-Scarce Reinforcement Learning

Xinjie Liu, Ruihan Zhao, Anirban Chaudhuri, Cyrus Neary, Ufuk Topcu, David Fridovich-Keil

arXiv 2610.02505首次发表:更新:

发表机构

University of Texas at Austin; University of British Columbia(德克萨斯大学奥斯汀分校; 不列颠哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MFPG-PPO,用低保真数据构建控制变量,稳定数据稀缺场景下的策略梯度强化学习,在模拟和真实机器人上均优于仅用高保真数据的PPO。

AI 中文摘要

针对在线策略强化学习(RL)的策略梯度方法,在昂贵且稀缺的目标域数据导致梯度估计噪声较大时,可能会变得不稳定。我们通过用丰富、廉价但有偏差的低保真(LF)数据(例如来自简化模拟器)补充有限的高保真(HF)目标域数据来应对这一挑战。大多数现有方法直接优化基于LF数据的有偏目标。相比之下,最近引入的多保真策略梯度(MFPG)框架仅使用LF数据来构建控制变量,以在不让策略梯度估计器产生偏差的情况下减少方差并提高HF数据效率。然而,已发表的MFPG工作仅限于在小型模拟任务上使用REINFORCE。我们将MFPG发展为适用于现代actor-critic学习,在GPU并行模拟和物理机器人上实现。我们的分析和实验表明,对近端策略优化(PPO)的朴素扩展可能会丢失跨保真相关性或使方差膨胀。我们的MFPG-PPO通过重新设计采样、优势估计和控制变量构建以保留跨保真相关性,并通过监控估计器不确定性以防止方差膨胀,从而解决了这些失败。我们还引入了预算感知的MFPG-PPO,以在高低保真数据源之间分配固定的采样预算。在模拟机器人运动任务中,跨越不同的LF到HF迁移难度和HF数据预算,MFPG-PPO在几乎所有设置中都优于仅使用HF数据训练的PPO,并且在最困难的任务上,在最小的HF预算下,始终匹配使用16倍更多HF数据训练的PPO的性能。相比之下,大多数使用LF数据的基线仅在直接LF到HF迁移成功时表现良好。MFPG-PPO使得在物理Franka机械臂上仅使用每次更新4个真实机器人episode且无需人工演示即可实现稳定学习。

英文摘要

Policy gradient methods for on-policy reinforcement learning (RL) can become unstable when expensive, scarce target-domain data yield noisy gradient estimates. We address this challenge by complementing limited high-fidelity (HF) target-domain data with abundant, cheap, but biased low-fidelity (LF) data, e.g., from a simplified simulator. Most existing methods directly optimize biased objectives based on LF data. In contrast, the recently introduced multi-fidelity policy gradient (MFPG) framework uses LF data solely to construct a control variate that reduces variance and improves HF data efficiency without biasing the policy gradient estimator. However, published work on MFPG is limited to REINFORCE on small-scale simulation tasks. We develop MFPG for modern actor-critic learning in GPU-parallel simulation and on a physical robot. Our analysis and experiments show that naive extensions to proximal policy optimization (PPO) can lose cross-fidelity correlation or inflate variance. Our MFPG-PPO addresses these failures by redesigning the sampling, advantage estimation, and control variate construction to preserve cross-fidelity correlation, and by monitoring estimator uncertainty to prevent variance inflation. We also introduce a budget-aware MFPG-PPO to divide a fixed sampling budget among high- and low-fidelity data sources. Across simulated robot locomotion tasks of varying LF-to-HF transfer difficulty and HF data budgets, MFPG-PPO improves upon PPO trained on HF data alone in nearly all settings, and consistently matches the performance of PPO trained with 16x more HF data on the hardest task at the smallest HF budgets. In contrast, most baselines that use LF data perform well only where direct LF-to-HF transfer succeeds. MFPG-PPO enables stable learning on a physical Franka arm using only 4 real-robot episodes per update and no human demonstrations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑