具有风险敏感性和评论家一致性正则化的鲁棒对抗强化学习
Robust Adversarial Reinforcement Learning with Risk Sensitivity and Critic Consistency Regularization
- Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对鲁棒对抗强化学习中的优化不稳定和价值估计退化问题,提出RACER框架,通过状态相关对抗目标和评论家一致性正则化,提升性能、鲁棒性与训练稳定性。
AI中文摘要:
强化学习(RL)在序列决策中取得了强劲的性能,但在动态不确定性和分布偏移下仍然脆弱。鲁棒对抗强化学习(RARL)通过最坏情况扰动来提高鲁棒性,但现有方法经常遭受不稳定的优化和退化的价值估计。特别是,过于激进的对抗者可能将智能体推向无信息的失败状态,而对抗性扰动会放大双评论家之间的分歧并引入有偏的价值目标。我们提出了一个统一框架,RACER(风险敏感的鲁棒对抗评论家一致性正则化强化学习),从风险敏感的角度重新审视对抗性强化学习。首先,我们引入了一个状态相关的对抗性目标,自适应地调节扰动强度,抑制有害干扰,同时保留信息丰富的探索。其次,我们提出了评论家一致性正则化,以减少Q值估计器之间的分歧并稳定学习。在具有挑战性的连续控制基准上的综合实验表明,RACER在强鲁棒RL基线上持续提高了性能、鲁棒性和训练稳定性。
英文摘要:
Reinforcement learning (RL) achieves strong performance in sequential decision-making but remains brittle under dynamic uncertainty and distributional shifts. Robust Adversarial Reinforcement Learning (RARL) improves robustness via worst-case perturbations, but existing approaches frequently suffer from unstable optimization and degraded value estimation. In particular, overly aggressive adversaries can drive the agent toward uninformative failure states, while adversarial perturbations amplify disagreement between double critics and introduce biased value targets. We propose a unified framework, RACER (Risk-sensitive robust Adversarial critic ConsistEncy-regularized Reinforcement learning), that revisits adversarial RL from a risk-sensitive perspective. First, we introduce a state-dependent adversarial objective that adaptively regulates perturbation strength, suppressing harmful disturbances while preserving informative exploration. Second, we propose critic consistency regularization to reduce disagreement between Q-value estimators and stabilize learning. Comprehensive experiments on challenging continuous control benchmarks demonstrate that RACER consistently improves performance, robustness, and training stability over strong robust RL baselines.