arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过对抗重要性采样的鲁棒策略优化

Robust Policy Optimization via Adversarial Importance Sampling

Amine Andam, Jamal Bentahar, Mustapha Hedabou

arXiv 2609.13044首次发表:更新:

发表机构

Mohammed VI Polytechnic University; Khalifa University; Concordia University(穆罕默德六世理工大学; 哈利法大学; 康考迪亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Advis方法,通过对抗重要性采样优化鲁棒策略,并引入advrl库和更全面的评估,在连续控制任务中验证了有效性。

AI 中文摘要

在保护深度强化学习(DRL)策略免受输入扰动方面已取得重大进展。开发鲁棒的DRL涉及三个主要阶段:算法设计、实现和评估。在这项工作中,我们识别并解决了每个阶段的一个关键局限性。首先,我们引入了对抗重要性采样(Advis),一种利用标准训练轨迹上的重要性采样来估计和优化可验证的最坏情况回报的方法。Advis满足先前工作未同时实现的三个理想标准:它不需要额外的环境交互,不需要辅助网络,并且能够捕捉长期鲁棒性。其次,我们引入了advrl,一个模块化的PyTorch库,提供现有鲁棒性方法和对抗攻击的干净、单文件实现,促进快速原型设计,并实现可复现和可追踪的评估。第三,我们重新审视了在学习的对抗者下的评估,并表明最优的对抗超参数不会跨智能体转移,这可能导致在使用有限的攻击者配置集时高估鲁棒性。因此,我们针对大量多样的攻击者评估策略,使用的配置数量比先前工作多6-14倍。最后,我们在连续控制环境中评估了我们的方法,展示了其相对于现有基线的有效性。代码可在以下网址获取:此https URL

英文摘要

Significant progress has been made in safeguarding deep reinforcement learning (DRL) policies against input perturbations. Developing robust DRL involves three main stages: algorithm design, implementation, and evaluation. In this work, we identify and address a key limitation at each stage. First, we introduce Adversarial Importance Sampling (Advis), a method that uses importance sampling over trajectories from standard training to estimate and optimize verifiable worst-case returns. Advis satisfies three desirable criteria not jointly achieved by prior work: it requires no additional environment interactions, no auxiliary networks, and captures long-term robustness. Second, we introduce advrl, a modular PyTorch library that provides clean, single-file implementations of existing robustness methods and adversarial attacks, facilitating rapid prototyping and enabling reproducible and traceable evaluations. Third, we revisit evaluation under learned adversaries and show that optimal adversarial hyperparameters do not transfer across agents, which can lead to an overestimation of robustness when using a limited set of attacker configurations. Accordingly, we evaluate policies against a large and diverse set of attackers, using 6-14x more configurations than prior work. Finally, we evaluate our approach on continuous control environments, demonstrating its effectiveness relative to existing baselines. The code is available at: https://github.com/AmineAndam04/advrl

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑