SafeBranch:面向具身智能体的分支对安全对齐方法
SafeBranch: Branch-Pair Safety Alignment for Embodied Agents
- Dongguk University(东国大学)
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对具身智能体的交互式安全问题,提出SafeBranch框架,通过自身不安全回滚生成的分支对对齐执行器安全性,在多基准及分布外变体上实现安全成功数约为未训练基线十倍的效果,且不牺牲任务成功。
AI中文摘要:
基于视觉语言模型的具身智能体可完成指令任务,但过程中常违反安全约束,该问题近期被定义为交互式安全。训练此类智能体安全行动难度大,因安全与任务成功是不同目标,且安全仅出现在轨迹中少数安全关键步骤。标准监督方法不足:模仿安全轨迹仅教授行为却不解释安全原因,对比任意安全与不安全轨迹会将安全信号与无关差异混合。我们提出SafeBranch框架,通过智能体自身不安全回滚产生的分支对,对齐具身执行器的安全性。SafeBranch将每个不安全回滚回退至导致违规的安全关键步骤,查询智能体的安全替代方案,将原始动作与该替代方案配对,使两个分支仅在该步骤存在差异。训练后的执行器在部署时无需评判器即可安全行动。在IS-Bench、SafetyALFRED及包含未见任务与物体的分布外变体上,该方法可靠处理安全问题且不牺牲任务成功,在未见物体变体上的安全成功数约为未训练基线的十倍。
英文摘要:
Vision-language-model-based embodied agents can complete instructed tasks but often violate safety constraints in the process, a problem recently framed as interactive safety. Training such agents to act safely is difficult, since safety and task success are distinct objectives, and safety arises only at a small number of safety-critical steps within a trajectory. Standard supervision is insufficient: imitating safe trajectories teaches behavior without explaining why it is safe, and contrasting arbitrary safe and unsafe trajectories mixes the safety signal with unrelated differences. We propose SafeBranch, a framework that aligns an embodied actor on safety through branch pairs constructed from the actor's own unsafe rollouts via environment rollback. SafeBranch rolls each unsafe rollout back to the safety-critical step that caused the violation, queries the actor for a safe alternative, and pairs the original action with the alternative so that the two branches differ only at that step. The trained actor acts safely at deployment with no critic in the loop. On IS-Bench, SafetyALFRED, and out-of-distribution variants with unseen tasks and objects, it handles safety reliably without sacrificing task success, achieving roughly ten times more safe successes than the untrained baseline on the unseen-object variant.