arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17386cs.RO

MANIGUARD:用于基于规范的机器人操作安全评估与改进的基准及数据集套件

MANIGUARD: A Benchmark and Data Suite for Specification-Grounded Safety Evaluation and Improvement of Robotic Manipulation

  • Northwestern University(西北大学)
  • Stanford University(斯坦福大学)
  • William & Mary(威廉玛丽学院)

机构由 AI 辅助整理,请以论文原文为准。

Yiyan Peng, Philip Wang, Simon Sinong Zhan, Yiqi Lyu, Zhenyang Ni, Jixin Yan, Fiorelli Wong, Ruochen Jiao, Hang Yin, Xinyu Cao, Huajie Shao, Manling Li, Ruohan Zhang, Qi Zhu

AI总结:

本研究提出ManiGuard基准与数据集套件,针对机器人操作基础模型策略,通过含23000余次滚动的实验证实安全需独立于任务成功评估,微调可提升安全表现但仍存差距。

AI中文摘要:

机器人操作的基础模型策略在任务成功率上进展迅速,但针对其是否安全完成任务的严格评估仍存在不足。我们推出ManiGuard,这是一个用于评估和改进基础模型操作安全性的基于规范的框架,包含ManiGuard-Bench任务套件及配套的带安全标注的轨迹生成流水线。ManiGuard-Bench将6类接触丰富的家庭任务按技能×约束分类法组织为200个锁定基础任务,安全规范独立于任务成功进行指定。每个任务在1种分布内扰动和4种单轴分布外扰动下进行评估,且安全规范保持固定,共得到1000个锁定场景。所有滚动执行均由基于LTL_f的自动机监控器在物理基础谓词上进行运行时检查,而非基于学习分类器或LLM评判,该评估在仿真环境和真实Franka平台上开展。该流水线将自动运动规划生成器与人类遥操作配对,由相同的逐步骤监控器进行标注,直接支持安全感知的微调;我们发布了8000个带安全标注的演示,每个基础任务对应40个。通过对超过23000次滚动执行评估零样本和微调后的视觉语言动作策略(VLAs),我们发现:(i)安全必须独立于任务成功进行评估,因为6%-21%的成功滚动执行违反了规范;(ii)在我们的套件上进行微调可将安全任务完成率从接近零提升至7.5%-29.8%,且“执行且安全”的行为比例从16%-40%提升至51%-72%;但(iii)仍存在无法通过扩大演示规模消除的差距,21%-42%的已执行滚动执行仍存在违规,6类任务中有2类对所有策略的安全完成率均低于2%,且这些失败在分布偏移和硬件上依然存在。

英文摘要:

Foundation-model policies for robotic manipulation are advancing rapidly on task success, but rigorous evaluation of whether they succeed safely is still lacking. We introduce ManiGuard, a specification-grounded framework for evaluating and improving the safety of foundation-model manipulation, comprising the ManiGuard-Bench task suite and a paired safety-annotated trajectory-generation pipeline. ManiGuard-Bench organizes six contact-rich household task families into 200 locked base tasks along a skill $\times$ constraint taxonomy, with safety specified independently of task success. Each task is evaluated under one in-distribution and four single-axis out-of-distribution perturbations that hold the safety specification fixed, giving 1,000 locked scenarios. Every rollout is runtime-checked by LTL$_f$-grounded automaton monitors over physics-grounded predicates rather than learned classifiers or LLM judges, in simulation and on a physical Franka platform. The pipeline pairs an automated motion-planning generator with human teleoperation, annotated by the same per-step monitor, and directly supports safety-aware fine-tuning; we release 8,000 safety-annotated demonstrations, 40 per base task. Benchmarking zero-shot and fine-tuned VLAs across more than 23,000 rollouts, we find: (i) safety must be evaluated independently of task success, as 6-21% of successful rollouts violate the specification; (ii) fine-tuning on our suite raises safe task completion from near zero to 7.5-29.8% and engaged-and-safe behavior from 16-40% to 51-72%; but (iii) a gap remains that scaling demonstrations does not close, with 21-42% of engaged rollouts still violating, two of six families below 2% safe success for every policy, and these failures persisting under distribution shift and on hardware.

↑