MANIGUARD:用于基于规范的机器人操作安全评估与改进的基准及数据集套件
MANIGUARD: A Benchmark and Data Suite for Specification-Grounded Safety Evaluation and Improvement of Robotic Manipulation
- Northwestern University(西北大学)
- Stanford University(斯坦福大学)
- William & Mary(威廉玛丽学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出ManiGuard基准与数据集套件,针对机器人操作基础模型策略,通过含23000余次滚动的实验证实安全需独立于任务成功评估,微调可提升安全表现但仍存差距。
AI中文摘要:
机器人操作的基础模型策略在任务成功率上进展迅速,但针对其是否安全完成任务的严格评估仍存在不足。我们推出ManiGuard,这是一个用于评估和改进基础模型操作安全性的基于规范的框架,包含ManiGuard-Bench任务套件及配套的带安全标注的轨迹生成流水线。ManiGuard-Bench将6类接触丰富的家庭任务按技能×约束分类法组织为200个锁定基础任务,安全规范独立于任务成功进行指定。每个任务在1种分布内扰动和4种单轴分布外扰动下进行评估,且安全规范保持固定,共得到1000个锁定场景。所有滚动执行均由基于LTL_f的自动机监控器在物理基础谓词上进行运行时检查,而非基于学习分类器或LLM评判,该评估在仿真环境和真实Franka平台上开展。该流水线将自动运动规划生成器与人类遥操作配对,由相同的逐步骤监控器进行标注,直接支持安全感知的微调;我们发布了8000个带安全标注的演示,每个基础任务对应40个。通过对超过23000次滚动执行评估零样本和微调后的视觉语言动作策略(VLAs),我们发现:(i)安全必须独立于任务成功进行评估,因为6%-21%的成功滚动执行违反了规范;(ii)在我们的套件上进行微调可将安全任务完成率从接近零提升至7.5%-29.8%,且“执行且安全”的行为比例从16%-40%提升至51%-72%;但(iii)仍存在无法通过扩大演示规模消除的差距,21%-42%的已执行滚动执行仍存在违规,6类任务中有2类对所有策略的安全完成率均低于2%,且这些失败在分布偏移和硬件上依然存在。
英文摘要:
Foundation-model policies for robotic manipulation are advancing rapidly on task success, but rigorous evaluation of whether they succeed safely is still lacking. We introduce ManiGuard, a specification-grounded framework for evaluating and improving the safety of foundation-model manipulation, comprising the ManiGuard-Bench task suite and a paired safety-annotated trajectory-generation pipeline. ManiGuard-Bench organizes six contact-rich household task families into 200 locked base tasks along a skill $\times$ constraint taxonomy, with safety specified independently of task success. Each task is evaluated under one in-distribution and four single-axis out-of-distribution perturbations that hold the safety specification fixed, giving 1,000 locked scenarios. Every rollout is runtime-checked by LTL$_f$-grounded automaton monitors over physics-grounded predicates rather than learned classifiers or LLM judges, in simulation and on a physical Franka platform. The pipeline pairs an automated motion-planning generator with human teleoperation, annotated by the same per-step monitor, and directly supports safety-aware fine-tuning; we release 8,000 safety-annotated demonstrations, 40 per base task. Benchmarking zero-shot and fine-tuned VLAs across more than 23,000 rollouts, we find: (i) safety must be evaluated independently of task success, as 6-21% of successful rollouts violate the specification; (ii) fine-tuning on our suite raises safe task completion from near zero to 7.5-29.8% and engaged-and-safe behavior from 16-40% to 51-72%; but (iii) a gap remains that scaling demonstrations does not close, with 21-42% of engaged rollouts still violating, two of six families below 2% safe success for every policy, and these failures persisting under distribution shift and on hardware.