arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13050cs.LG

基于正则化的鲁棒强化学习:一个统一且受约束的视角

A Unified and Constrained View of Regularization-Based Robust Reinforcement Learning

Amine Andam, Jamal Bentahar, Mustapha Hedabou

首次发表
浏览论文内容

中文总结 AI 辅助

本文统一了基于正则化的鲁棒强化学习方法,推导出性能差距上界,并将鲁棒训练转化为约束优化问题,通过联合更新拉格朗日乘子自动调整正则化权重,在连续控制任务上验证了理论分析。

中文摘要 AI 辅助

基于正则化的方法已成为训练深度强化学习策略以应对对抗性输入扰动的一种标准途径。在本文中,我们通过推导名义策略与最坏情况策略之间性能差距的新上界来统一这些方法。每个上界均表示为现有的正则化目标加上名义策略与最坏情况策略之间的KL散度惩罚项,这进一步解释了为何在实践中添加KL惩罚能够提升鲁棒性。基于这些界,我们将鲁棒训练表述为一个约束优化问题,并表明现有方法对应于固定拉格朗日乘子的特殊情况。我们转而将乘子与策略联合更新,以自动调整正则化权重。最后,我们在多个连续控制任务上进行了广泛的对抗性评估,以验证我们的理论分析。

英文摘要

Regularization-based methods have become a standard approach for training Deep Reinforcement Learning policies against adversarial input perturbations. In this paper, we unify these methods by deriving new upper bounds on the performance gap between the nominal and worst-case policies. Each upper bound is expressed as an existing regularization objective plus a KL-divergence penalty between the nominal and worst-case policies, which further explains why adding a KL penalty improves robustness in practice. Building on these bounds, we formulate robust training as a constrained optimization problem, showing that existing methods correspond to the special case of a fixed Lagrange multiplier. We instead update the multiplier jointly with the policy to automatically tune the regularization weight. Finally, we conduct extensive adversarial evaluations across several continuous control tasks to validate our theoretical analysis.

发表机构

  • Mohammed VI Polytechnic University(穆罕默德六世理工大学)
  • Khalifa University(哈利法大学)
  • Concordia University(康考迪亚大学)

机构由 AI 辅助整理,请以论文原文为准。

↑