发表机构
Indian Institute of Science (IISc); Microsoft Research(印度科学理工学院; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ShieldVLA提出基于HJ可达性的安全对齐微调框架,通过无模型价值函数估计安全区域并分离奖励与恢复,在五个基准上平均降低57%安全成本并提升成功率0.13。
AI 中文摘要
视觉-语言-动作(VLA)模型在机器人操作和导航中展现出强大的泛化能力,但现有的微调方法提供的安全保证有限。当前方法主要依赖拉格朗日优化,通过对期望累积成本施加软惩罚来强制安全,往往导致残余约束违反或过度保守的行为。此外,由于缺乏密集的逐步骤安全标注,在视觉领域中学习安全具有挑战性。我们提出ShieldVLA,一种基于Hamilton-Jacobi(HJ)可达性的VLA模型安全对齐微调框架。ShieldVLA直接从视觉观测中学习HJ可达性价值函数的无模型近似,以估计安全操作区域。学习到的安全评论家通过将可行区域内的奖励最大化与不安全状态附近的恢复分离来门控策略优化,避免了持续的奖励-成本权衡。为了在视觉环境中实现可扩展的监督,我们引入了基于规则(rubric)的VLM安全评分,将语义安全反馈转换为结构化的评论家目标,无需手动成本标签。在跨越多个VLA骨干的五个导航和操作基准上,ShieldVLA平均将累积安全成本降低了57%,并将任务成功率相对于SafeVLA提高了+0.13。
英文摘要
Vision-Language-Action (VLA) models demonstrate strong generalization in robotic manipulation and navigation, but existing fine-tuning methods provide limited safety guarantees. Current approaches primarily rely on Lagrangian optimization that enforces safety through soft penalties on expected cumulative cost, often resulting in residual constraint violations or overly conservative behavior. Moreover, learning safety in visual domains is challenging due to the absence of dense per-step safety annotations. We propose ShieldVLA, a safety-aligned fine-tuning framework for VLA models based on Hamilton-Jacobi (HJ) reachability. ShieldVLA learns a model-free approximation of the HJ reachability value function directly from visual observations to estimate the safe operating region. The learned safety critic gates policy optimization by separating reward maximization within feasible regions from recovery near unsafe states, avoiding persistent reward-cost trade-offs. To enable scalable supervision in visual environments, we introduce rubric-based VLM safety scores that convert semantic safety feedback into structured critic targets without requiring manual cost labels. Across five navigation and manipulation benchmarks spanning multiple VLA backbones, ShieldVLA reduces cumulative safety cost by 57% on average and improves task success rate by +0.13 over SafeVLA.
Comments24 pages, 3 figures, 14 tables