arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10866cs.LGmath.OC

对抗状态扰动下风险敏感强化学习的下界认证

Certifying Lower Bounds for Risk-Sensitive Reinforcement Learning under Adversarial State Perturbations

Tong Li, Saunak Kumar Panda, Yisha Xiang

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对风险敏感强化学习,在对抗状态扰动下通过凸优化和phi散度松弛推导认证下界,实验表明风险厌恶训练提升下界,但过度风险厌恶会降低性能。

中文摘要 AI 辅助

部署在真实环境中的强化学习(RL)智能体常常容易受到状态观测中对抗性扰动的影响,这在安全关键应用中带来了风险。认证方法可以通过提供期望累积奖励的下界来增强对对抗性扰动的鲁棒性。然而,现有的认证方法主要关注风险中性的目标。在本文中,我们将认证方法扩展到风险敏感目标,通过建立$l_{p}$范数有界状态对抗扰动($1\leq p <\infty$)下累积奖励指数效用的下界。通过引入扰动集的$\phi$-散度松弛,我们将风险敏感认证问题表述为凸优化,并推导其对偶形式,以获得认证下界的可处理近似。我们进一步提出了一种经验方法,通过独立于评估时使用的风险水平选择训练风险厌恶参数$\beta$来改进认证下界。在OpenAI Gym环境和机器更换问题上的实验表明,与风险中性训练相比,风险厌恶训练通常产生具有更高认证下界的策略,特别是在较大扰动预算下。此外,在风险中性和风险厌恶评估设置下,训练期间增加风险厌恶会导致非单调的认证性能,认证下界最初改善但最终因过于保守的策略而下降。

英文摘要

Reinforcement learning (RL) agents deployed in real-world environments are often vulnerable to adversarial perturbations in state observations, creating risks in safety-critical applications. Certification methods can improve robustness against adversarial perturbations by providing lower bounds on expected cumulative rewards. Existing certification methods, however, mainly focus on risk-neutral objectives. In this paper, we extend certification methods to risk-sensitive objectives by establishing lower bounds on the exponential utility of cumulative rewards under $l_{p}$-norm-bounded state adversarial perturbations ($1\leq p <\infty$). By introducing a $ϕ$-divergence relaxation of the perturbation set, we formulate the risk-sensitive certification problem as a convex optimization and derive its dual to obtain a tractable approximation of the certified lower bound. We further propose an empirical method that improves certified lower bounds by selecting the training risk-aversion parameter $β$ independently of the risk level used during evaluation. Experiments on both OpenAI Gym environments and a machine replacement problem show that, compared to risk-neutral training, risk-averse training generally yields policies with higher certified lower bounds, particularly under larger perturbation budgets. Moreover, under both risk-neutral and risk-averse evaluation settings, increasing risk aversion during training leads to non-monotonic certification performance, where certified lower bounds initially improve but eventually decrease due to overly conservative policies.

发表机构

  • University of Houston(休斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑