发表机构
Shanghai Jiao Tong University; Shanghai Institute of Innovation(上海交通大学; 上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ActSafeGuard提出一种可微分且训练对齐的安全防护层,通过解析射线缩放算子将硬约束融入流匹配策略学习,实现100%逐步安全率并保持或提升任务成功率。
AI 中文摘要
视觉-语言-动作(VLA)模型和世界-动作模型(WAMs)在通用机器人操作中展现出强大的能力,然而它们生成的动作可能违反硬物理约束,从而在部署时存在不安全或不可行的问题。现有的安全方法要么优化统计安全目标而不提供确定性的逐步保证,要么仅在推理期间纠正不安全动作,导致策略训练与执行之间的不匹配。我们引入了ActSafeGuard,一种用于基于流匹配策略的可微分且训练对齐的安全防护层。ActSafeGuard将硬动作可行性整合到策略学习中,而不仅仅将安全视为推理时的外部组件。通过解析射线缩放算子设计,ActSafeGuard能够提供边界感知梯度,引导模型自然学习约束流形。在多个标准基础骨干网络(π₀.₅和Fast-WAM)上的大量实验表明,ActSafeGuard始终实现100%的逐步安全率,同时完全保持甚至提升任务成功率,为安全具身AI部署提供了一种可扩展且最小侵入性的解决方案。
英文摘要
Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasible for deployment. Existing safety approaches either optimize statistical safety objectives without deterministic per-step guarantees or correct unsafe actions only during inference, creating a mismatch between policy training and execution. We introduce ActSafeGuard, a differentiable and training-aligned safeguard layer for flow-matching based policies. ActSafeGuard integrates hard action feasibility into policy learning, not merely treating safety as an inference-time external component. Through an analytical ray-scaling operator design, ActSafeGuard enables boundary-aware gradients to guide the model to naturally learn constrained manifolds. Extensive experiments on multiple standard foundation backbones ($π_{0.5}$ and Fast-WAM) across various tasks demonstrate that ActSafeGuard consistently achieves a $100\%$ step safety rate while fully preserving or even boosting task success rates, providing a scalable and minimally invasive solution for safe embodied AI deployment.