arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21223cs.RO

SafeStage:在视觉-语言条件机器人操作之前、期间和之后评估安全性

SafeStage: Evaluating Safety Before, During, and After Vision-Language-Conditioned Robot Manipulation

Jinzhu Luo, Qi Zhang, Wei Wang, Wei Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

SafeStage提出一个生命周期分阶段的基准,通过97个风险场景评估视觉-语言条件机器人操作在执行前、中、后的安全性,发现任务成功常伴随安全违规,并提供统一诊断平台。

中文摘要 AI 辅助

视觉-语言条件机器人策略整合了感知、语言理解和控制,用于通用操作任务。然而,现有评估往往侧重于任务成功率、孤立的物理约束、语义拒绝或已实现的物理损伤,对闭环操作中安全失效的具体环节提供的洞察有限。我们提出了SafeStage,一个生命周期结构化的基准,用于评估任务执行前、执行中和执行后的操作安全性。SafeStage包含97个专门构建的风险场景,分为三个阶段。初始状态危险(Initial-State Hazards)捕捉在操作目标之前必须解决的安全相关关系。执行时安全(Execution-Time Safety)评估执行过程中的不安全接触、轨迹、区域进入和对象交互。最终状态危险(Final-State Hazards)捕捉名义任务完成后遗留的不稳定或其他不安全条件。该基准使用基于事件和基于状态的检查来评估实际交互,并独立于阶段特定的安全结果报告原生任务成功率。我们在统一的闭环协议下评估了具有代表性的直接动作视觉-语言-动作(VLA)策略和基于世界模型的策略。我们的结果表明,名义任务完成经常与安全违规共存,且不同策略在三个阶段中表现出不同的失败特征。通过将任务成功与安全分离,并定位违规发生的时间,SafeStage为评估和改进视觉-语言条件机器人操作策略提供了一个统一的诊断测试平台。

英文摘要

Vision-language-conditioned robot policies integrate perception, language understanding, and control for general-purpose manipulation. However, existing evaluations often focus on task success, isolated physical constraints, semantic refusal, or realized physical damage, providing limited insight into where safety fails during closed-loop manipulation. We introduce SafeStage, a lifecycle-structured benchmark for evaluating manipulation safety before, during, and after task execution. SafeStage contains 97 purpose-built risk scenarios organized into three stages. Initial-State Hazards captures safety-relevant relations that must be resolved before manipulating the target. Execution-Time Safety evaluates unsafe contacts, trajectories, region entries, and object interactions during execution. Final-State Hazards capture unstable or otherwise unsafe conditions remaining after nominal task completion. The benchmark evaluates realized interactions using event-based and state-based checks and reports native task success independently from stage-specific safety outcomes. We evaluate representative direct-action Vision-Language-Action (VLA) policies and policies with world-model-based policies under a common closed-loop protocol. Our results demonstrate that nominal task completion frequently coexists with safety violations and that different policies exhibit distinct failure profiles across the three stages. By separating task success from safety and localizing when violations occur, SafeStage provides a unified diagnostic testbed for evaluating and improving vision-language-conditioned robot manipulation policies.

发表机构

  • Worcester Polytechnic Institute(伍斯特理工学院)
  • Futurewei Technologies Inc.(未来华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑