arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

平均安全,尾部不安全:剧集成本尾部何时可控?

Safe on Average, Unsafe in the Tail: When Is the Episodic-Cost Tail Controllable?

Samuel Tetteh, Cody Fleming

arXiv 2610.09508首次发表:更新:

发表机构

Iowa State University(爱荷华州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究用CVaR_0.1衡量剧集成本尾部,识别平均安全但尾部不安全的策略,并探究在保持回报的同时控制尾部违规的可行性。

AI 中文摘要

安全强化学习旨在寻找在满足累积成本约束的同时最大化回报的策略。大多数方法将约束施加在期望剧集成本上。因此,标准评估报告平均剧集成本,而不描述成本在剧集间的分布。满足平均成本准则的策略可能在其最差剧集中仍然不安全。平均成本报告既不能识别这种尾部违规,也不能显示在保持回报的同时能否将其控制在预算内。在本工作中,我们使用 $\mathrm{CVaR}_{0.1}$(最差 $10\\%$ 剧集的平均成本)来衡量剧集成本尾部。当 $\mathrm{CVaR}_{0.1}$ 在安全预算内时,我们将策略分类为尾部安全。这使我们首先能够识别那些平均安全但尾部不安全的策略,然后研究在保持回报的同时其尾部违规是否可控。为识别尾部不安全的策略,我们在三个 Safety-Gymnasium 导航任务上评估了五种标准算法。随后,我们在密集危险导航任务上考察了四类约束族,并在四个导航和四个运动任务上评估尾部控制能力。

英文摘要

Safe reinforcement learning seeks policies that maximize return while satisfying constraints on cumulative cost. Most methods impose these constraints on expected episodic cost. Consequently, standard evaluations report mean episodic cost without characterizing how cost is distributed across episodes. A policy that satisfies the mean-cost criterion may therefore remain unsafe in its worst episodes. Mean-cost reporting neither identifies this tail violation nor shows whether it can be brought within budget while preserving return. In this work, we measure the episodic-cost tail using $\mathrm{CVaR}_{0.1}$, the average cost of the worst $10\%$ of episodes. We classify a policy as tail-safe when $\mathrm{CVaR}_{0.1}$ is within the safety budget. This allows us first to identify policies that are safe on average but unsafe in the tail and then to study whether their tail violations can be controlled while preserving return. To identify tail-unsafe policies, we evaluate five standard algorithms on three Safety-Gymnasium navigation tasks. We then examine four constraint families on dense-hazard navigation and assess tail control across four navigation and four locomotion tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑