arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12432cs.ROcs.LGcs.SYeess.SY

FAITH:面向高维系统的可行性感知安全过滤强化学习

FAITH: Feasibility-Aware Safety-Filtered RL for High-Dimensional Systems

Songyuan Zhang, Baljeet Singh, Sarthak Ranjeet Kaingade, Chuchu Fan, Bryan Trinh

首次发表
浏览论文内容

中文总结 AI 辅助

FAITH 是一种无模型可行性感知安全过滤强化学习框架,可近似最优安全值并摊销最小干预过滤,在双积分器、Safety Gym 及 Unitree G1 人形机器人任务中实现了高安全率与任务性能的平衡。

中文摘要 AI 辅助

安全强化学习通常将安全性和任务性能置于同一策略目标中,这可能会引入相互竞争的更新。安全过滤器在动作执行阶段将二者分离,但经典设计需要解析型安全函数和动力学模型,且标准的最小干预过滤器因仅最小化瞬时动作偏差,对长 horizon 任务回报存在短视问题;当不存在安全动作时,硬投影也无定义。我们提出 FAITH,这是一个可行性感知的无模型框架,可近似最优状态-动作安全值,并通过前馈网络摊销最小干预过滤。任务策略通过过滤后的动力学优化任务回报,在任务策略更新中无需竞争安全项即可恢复可行的约束问题。当无动作满足学习到的安全条件时,同一过滤器会趋近于具有最小预测峰值危害的动作。在双积分器示例和 Safety Gym 环境中,FAITH 在无可行起始违规的方法中实现了最高回报,并匹配了不可行起始的最低危害;在 29 自由度人形机器人上,它在 Walking-Avoid 任务中达到 99.95% 的安全率,同时保留了 97% 的未过滤回报,在 Push-Avoid 任务中通过学习牺牲平衡并远离受保护区域获得了最高测量安全率,相同策略还在现实世界的 Unitree G1 人形机器人上得到验证。

英文摘要

Safe reinforcement learning commonly places safety and task performance in the same policy objective, where they can introduce competing updates. Safety filters separate them at action execution, but classical designs require an analytic safety function and dynamics model, and standard minimal-intervention filters are myopic to long-horizon task return because they minimize only instantaneous action deviation. Hard projections are also undefined when no safe action exists. We present FAITH, a feasibility-aware, model-free framework that approximates the optimal state-action safety value and amortizes minimal-intervention filtering with a feedforward network. The task policy optimizes the task return through the filtered dynamics, which recovers the feasible constrained problem without a competing safety term in the task-policy update. When no action satisfies the learned safety condition, the same filter approaches the action with minimum predicted peak harm. On a double integrator example and a Safety Gym environment, FAITH achieves the highest return among methods with no feasible-start violations and matches the lowest harm from infeasible starts. On a 29-DoF humanoid, it reaches a 99.95% safety rate while retaining 97% of the unfiltered return in Walking-Avoid, and obtains the highest measured safety rate in Push-Avoid by learning to sacrifice balancing and fall away from the protected region. The same policies are also demonstrated on a real-world Unitree G1 humanoid.

发表机构

  • Massachusetts Institute of Technology(麻省理工学院)
  • Amazon(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑