发表机构
Politecnico di Milano; Università degli Studi di Milano(米兰理工大学; 米兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在多重重要性加权框架下提出wPPO-U和wPPO-BH两种PPO变体,通过重用近期窗口样本并推导策略改进下界,系统探究数据重用对提升样本效率与最终性能的时机与方式。
AI 中文摘要
在基于策略的深度强化学习方法中,近端策略优化(PPO)已成为事实上的标准,因为它在不同应用领域均展现出持续强劲的经验性能。然而,基于策略的方法本质上样本效率较低:在当前策略下收集的新鲜数据仅用于几次更新便被丢弃。基于离策略的方法通过经验回放避免了这种低效,实现了显著的样本效率提升,但代价是训练不稳定或需要大量调参。这促使了混合策略的兴起,即用离策略数据重用增强PPO。现有的PPO样本重用变体相较于原始PPO展示了改进的样本效率,但关于重用何时有帮助、在哪些场景下以及帮助程度如何的系统性研究仍然缺失。在本工作中,我们通过在多重重要性加权框架内实例化两个变体来研究PPO中样本重用的有效性。两者都保留了PPO的核心机制,仅重用来自最近迭代窗口的样本,从而将数据重用的效果与其他因素隔离。这两个变体分别称为wPPO-U和wPPO-BH,分别使用普通重要性权重或平衡启发式校正权重。对于两者,我们推导了策略改进下界,为其各自的损失提供了理论依据。我们利用它们来实证研究数据重用何时以及如何提高PPO在连续控制任务中的样本效率或最终性能。
英文摘要
Among on-policy deep reinforcement learning methods, Proximal Policy Optimization (PPO) has become the de facto standard, due to its consistently strong empirical performance across diverse application domains. However, on-policy methods are inherently sample inefficient: fresh data collected under the current policy is used for just a few updates before being discarded. Off-policy methods avoid this inefficiency via experience replay, achieving notable sample efficiency gains, but at the cost of training instabilities or extensive tuning. This motivated the rise of hybrid strategies that augment PPO with off-policy data reuse. Existing sample-reuse variants of PPO demonstrated improved sample efficiency over vanilla PPO, yet a systematic study of when reuse helps, in which scenarios, and to what extent remains missing. In this work, we study the effectiveness of sample reuse in PPO by instantiating two variants within a multiple importance weighting framework. Both retain the core PPO mechanics, reusing only samples from a window of recent iterations, thereby isolating the effect of data reuse from other factors. The variants, termed wPPO-U and wPPO-BH, employ vanilla importance weights or balance-heuristic-corrected ones, respectively. For both, we derive policy improvement lower bounds providing theoretical grounding for their respective losses. We use them to empirically study when and how data reuse improves sample efficiency or final performance of PPO across continuous control tasks.