面向移动危险物的 VLA 策略多链路安全过滤
Multi-Link Safety Filtering for VLA Policies Around Moving Hazards
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对 VLA 策略在移动危险物环境中的安全问题,提出免训练的多链路椭球防护罩与光流跟踪,显著降低碰撞率并提升安全成功率。
AI中文摘要:
视觉-语言-动作(VLA)策略在完成操作任务时可能会碰倒与其无关的物体,因此仅凭任务成功并不能证明该策略在杂乱环境中部署是安全的。我们研究如何在运行时保持预训练的 VLA 策略远离此类危险物而无需重新训练,这要求对机械臂的更多部分(而非仅末端执行器)进行防护,跟随危险物的移动,并与策略共享机载计算资源。我们的免训练防护罩用五个椭球体覆盖夹爪、腕部和前臂,并通过一个障碍程序对每个指令运动进行过滤,该程序基于重置时从 RGB-D 感知拟合的禁入椭球体。随后,稀疏光流将该椭球体的中心随危险物一起移动,无需重复检测或重新拟合。在六种模拟危险物运动条件下,防护罩将碰撞率从 65.62% 降至 27.27%,并将安全成功率(无碰撞的任务完成率)从 29.35% 提升至 50.43%。消融实验表明,防护手臂连杆可提供超出末端执行器防护的保护,而跟踪则恢复了在重置时危险物估计被冻结所失去的大部分保护。在异构边缘硬件上,五椭球体障碍程序在 CPU 上运行,第 99 百分位耗时 2.2 毫秒;精简视觉-语言前缀并减少流匹配步骤,将集成 GPU 上每次 π0.5 策略调用从 343 毫秒缩短至 177.3 毫秒。在物理 SO-101 机械臂上跨四个任务,防护的 16 次试验中手臂接触危险物 3 次,而未防护的 16 次试验中接触 11 次。项目页面:此 https URL
英文摘要:
A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy. Our training-free shield covers the gripper, wrist, and forearm with five ellipsoids and filters every commanded motion through one barrier program against a keep-out ellipsoid fitted from RGB-D perception at reset. Sparse optical flow then carries that ellipsoid's center along with the hazard, with no repeated detection or refitting. Over six simulated hazard-motion conditions, the shield lowers collision from $65.62\%$ to $27.27\%$ and raises safe-success, task completion without collision, from $29.35\%$ to $50.43\%$. Ablations show that guarding the arm links protects beyond end-effector shielding, and that tracking recovers most of the protection lost when the hazard estimate is frozen at reset. On heterogeneous edge hardware, the five-ellipsoid barrier runs on the CPU in $2.2$~ms at the 99th percentile, and trimming the vision--language prefix and taking fewer flow-matching steps shortens each $π_{0.5}$ policy call on the integrated GPU from $343$ to $177.3$~ms. On a physical SO-101 arm across four tasks, the arm touched the hazard in 3 of 16 shielded episodes versus 11 of 16 unshielded ones. Project page: https://yathag.github.io/multilink-safety-filter/