发表机构
School of Computer Science and Technology, Tongji University; Shanghai University of Engineering Science(同济大学计算机科学与技术学院; 上海工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CWAC框架,通过分布式评判者、协同加权机制及随机悲观值估计缓解离线强化学习的高估问题,可集成于SAC等算法,在多模拟任务中性能显著提升。
AI 中文摘要
用于连续控制的深度离线强化学习算法通常依赖神经值函数近似来指导策略改进。然而,时序差分(TD)学习会引入带噪目标,导致非平稳优化,而贪心策略更新会放大早期估计误差。这类误差的递归传播会导致演员-评判者方法中持续的高估偏差和训练稳定性下降。现有方法试图通过优先采样或修改值学习目标来缓解该问题,但往往过度强调因数据覆盖有限或自举误差导致的高不确定性转移,从而进一步放大问题。本文提出协同加权演员-评判者(CWAC),这是一个明确考虑值估计中预测不确定性的统一框架。CWAC采用分布式评判者对回报不确定性进行建模,并引入协同加权机制,该机制联合重加权TD误差和不确定性,从而能从可靠样本中进行稳健学习,同时抑制带噪更新。此外,我们通过从回报分布中采样纳入随机悲观值估计方案,有效缓解策略改进期间的误差传播。CWAC可无缝集成到现有离线算法框架(如SAC、TD3和DDPG)中,且仅需极少开销。实验结果表明,所提方法在多种模拟任务中显著提升了性能。我们的代码可在this https URL获取。
英文摘要
Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement. However, temporal-difference (TD) learning introduces noisy targets, resulting in non-stationary optimization, while greedy policy updates amplify early-stage estimation errors. The recursive propagation of such errors leads to persistent overestimation bias and degraded training stability in actor-critic methods. Existing approaches attempt to alleviate this issue via prioritized sampling or modified value learning objectives, but often overemphasize high-uncertainty transitions caused by limited data coverage or bootstrapping errors, thereby further amplifying bias.In this paper, we propose Collaborative Weighting Actor-Critic (CWAC), a unified framework that explicitly accounts for predictive uncertainty in value estimation. CWAC employs distributional critic to model return uncertainty and introduces a collaborative weighting mechanism that jointly reweights TD-errors and uncertainty, enabling robust learning from reliable samples while suppressing noisy updates. In addition, we incorporate a stochastic pessimistic value estimation scheme via sampling from the return distribution, which effectively mitigates error propagation during policy improvement. CWAC can be seamlessly integrated into existing off-policy algorithm frameworks such as SAC, TD3, and DDPG with minimal overhead. Empirical results demonstrate that our proposed method significantly enhances performance across a diverse range of simulated tasks. Our code is publicly available at https://anonymous.4open.science/r/CWAC-348E.