PlanGuard:具身智能体中多步计划安全性的防护栏
PlanGuard: A Guardrail for Multi-Step Plan Safety in Embodied Agents
浏览论文内容
中文总结 AI 辅助
PlanGuard提出首个执行前多步计划物理安全检测器,通过MSP-Safe数据集训练及STAC-OPD蒸馏方法,使紧凑模型PlanGuard-2B达到87.15%准确率和87.21%F1分数,有效检测整体计划风险。
中文摘要 AI 辅助
具身任务规划器可能生成多步计划,其子任务依赖关系及与环境交互在执行过程中会产生物理风险。然而,现有的安全防护措施忽视了此类组合风险,因为通用防护栏侧重于语义危害,而具身安全检测器则孤立地评估子任务。为填补这一空白,我们引入了PlanGuard,这是首个在执行前检测器,用于评估完整多步计划在当前环境中的物理安全性。为进行训练和评估,我们通过配对任务构建、使用多种规划器生成计划以及由三位评审员进行安全标注,构建了多步计划安全(MSP-Safe)数据集。在MSP-Safe上进行面向任务的SFT建立了基本的计划安全评估能力,但适用于实时部署的紧凑模型与更强但成本更高的大型模型之间仍存在显著差距。因此,我们提出了强教师自适应补偿用于在线策略蒸馏(STAC-OPD),该方法沿在线策略轨迹为紧凑模型提供自适应强教师监督。它结合了来自微调强教师的词元级分布转移与基于概率路由的序列级补偿,当学生倾向于参考安全决策时保留学生生成的目标,否则使用教师重建的目标。在所有测试子集上,PlanGuard-2B实现了平均87.15%的准确率和87.21%的F1分数,展示了在紧凑模型规模下有效的整体计划物理风险检测。代码和数据集将公开发布。
英文摘要
Embodied task planners may produce multi-step plans whose subtask dependencies and interactions with the environment create physical risks during execution. Yet existing safeguards overlook such compositional risks, as general-purpose guardrails focus on semantic harm and embodied safety detectors assess subtasks in isolation. To address this gap, we introduce PlanGuard, the first pre-execution detector that evaluates the physical safety of a complete multi-step plan in its current environment. For training and evaluation, we construct a Multi-Step Plan Safety (MSP-Safe) dataset through paired task construction, plan generation using diverse planners, and safety annotation by three judges. Task-oriented SFT on MSP-Safe establishes fundamental plan-safety assessment capabilities, yet a substantial gap remains between compact models suitable for real-time deployment and stronger but costlier large models. Accordingly, we propose Strong-Teacher Adaptive Compensation for On-Policy Distillation (STAC-OPD), which provides compact models with adaptive strong-teacher supervision along their on-policy trajectories. It combines token-level distribution transfer from a fine-tuned strong teacher with probability-routed sequence-level compensation, retaining student-generated targets when the student favors the reference safety decision and using teacher-reconstructed targets otherwise. Across all test subsets, PlanGuard-2B achieves average 87.15% ACC and 87.21% F1, demonstrating effective whole-plan physical-risk detection at compact model scale. Code and dataset will be publicly released.