发表机构
LMU Munich; TUM; Masaryk University; CU Boulder(慕尼黑大学; 慕尼黑工业大学; 马萨里克大学; 科罗拉多大学博尔德分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FoSeRL提出一种无需修改策略、基于可重置模拟器的框架,通过时间相关屏障条件认证强化学习策略在稀疏动作攻击下的回报损失上限,并在多个环境中验证了其有效性。
AI 中文摘要
即使少量的动作扰动也能显著降低已部署决策策略的性能。在随机环境中,认证由此产生的回报损失具有挑战性,因为在没有攻击的情况下回报也会变化。我们引入了FoSeRL,一个用于认证已部署的强化学习策略免受预先承诺的、时间上稀疏的动作攻击的框架。已部署的策略保持不变,无需平滑或重新训练。认证需要一个支持共享随机性和独立一步后继查询的可重置模拟器,但不需要解析动力学模型。FoSeRL认证,一个受攻击的回合相对于同一未受攻击的回合,其回报损失不超过规定数量,且至少以目标概率和用户指定的置信水平成立。两次运行共享初始状态和随机性,因此测量的损失反映的是攻击而非回合本身;将运行中的回报差距作为状态坐标,使其成为终端值,从而将轨迹级认证简化为终端安全性。对增强状态的时间相关屏障条件限制了终端失败概率:若精确满足,它们认证所有可接受的预先承诺攻击;若从采样轨迹中学习并在保留数据上验证,它们则在指定的攻击回合设置下认证相同的保证。在六个随机连续控制环境和三个强化学习策略家族(TD3、SAC和PPO)中,FoSeRL认证了非平凡的基数-幅度鲁棒性边界,实现了比策略平滑大得多的认证预算,并揭示了具有可比名义性能的策略之间显著的鲁棒性差异。
英文摘要
Even a few action perturbations can substantially degrade the performance of a deployed decision policy. Certifying the resulting return loss is challenging in stochastic environments, where returns vary even without an attack. We introduce FoSeRL, a framework for certifying deployed RL policies against precommitted, temporally sparse action attacks. The deployed policy is unchanged, with no smoothing or retraining. Certification requires a resettable simulator supporting shared randomness and independent one-step successor queries, but no analytical dynamics model. FoSeRL certifies that an attacked episode loses no more than a prescribed amount of return relative to the same episode unattacked, with at least a target probability and at a user-specified confidence level. Both runs share the initial state and randomness, so the measured loss reflects the attack, not the episode; carrying the running return gap as a state coordinate makes it the terminal value, reducing trajectory-level certification to terminal safety. Time-dependent barrier conditions on the augmented state bound the terminal failure probability: satisfied exactly, they certify every admissible precommitted attack; learned from sampled trajectories and verified on held-out data, they certify the same guarantee under a specified attack-episode setting. Across six stochastic continuous-control environments and three RL policy families (TD3, SAC, and PPO), FoSeRL certifies non-trivial cardinality--magnitude robustness frontiers, achieves substantially larger certified budgets than policy smoothing, and reveals marked robustness differences among policies with comparable nominal performance.