arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.20698cs.RO

SafeDojo:基于交互世界模型的安全强化学习用于视觉-语言-动作模型

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model

Kai Tang, Peidong Jia, Zhong Chu, Jixian Wu, Rui Ma, Jiajun Cao, Fangyuan Zhao, Sixiang Chen, Yichen Guo, Xiaowei Chi, Chun-Kai Fan, Kevin Zhang, Jinchang Xu, F… 展开作者

Kai Tang, Peidong Jia, Zhong Chu, Jixian Wu, Rui Ma, Jiajun Cao, Fangyuan Zhao, Sixiang Chen, Yichen Guo, Xiaowei Chi, Chun-Kai Fan, Kevin Zhang, Jinchang Xu, Fubing Yang, Weishi Mi, Xiaozhu Ju, Jian Tang, Shanghang Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

提出SafeDojo,首个基于模型的安全强化学习框架,通过交互式视频世界模型进行想象学习安全动作,结合解耦的任务奖励和安全代价信号,在SafeLIBERO和真实机器人上取得最佳安全成功率。

中文摘要 AI 辅助

安全控制是现实世界具身智能的先决条件,安全强化学习已成为一种有前景的范式。然而,现有安全强化学习方法要么需要昂贵的真实世界探索,要么依赖手工设计的安全函数。两者都无法扩展到部署在开放世界物理环境中的视觉-语言-动作模型。我们提出SafeDojo,首个基于模型的视觉-语言-动作策略安全强化学习框架,旨在通过基于世界模型的想象学习安全动作。具体地,SafeDojo在交互式视频世界模型之上执行在线强化学习。世界模型生成动作条件下的未来预测,从中定制的ResNet成功分类器从想象帧中估计每步任务进度,轻量级安全头从潜在上下文连同提出的动作块预测每步安全代价,从而同时评估任务执行和轨迹安全性。解耦的任务奖励和安全代价信号通过基于拉格朗日的约束GRPO目标进行平衡,实现在显式约束下任务成功和安全的协调改进。在SafeLIBERO上,SafeDojo在推理时安全、无模型RL和基于模型RL基线中取得了最佳的综合任务成功、安全成功和执行效率,在两个级别上均取得最佳平均安全成功率,并在Level I上比最强基线提高8.25个百分点。真实Franka部署进一步展示了在五个任务上最佳的平均任务和安全成功率。我们的结果将基于世界模型的安全强化学习定位为通向安全具身智能的可扩展和可泛化路径。

英文摘要

Safe control is a prerequisite for real-world embodied intelligence, for which safe reinforcement learning has emerged as a promising paradigm. However, existing safe reinforcement learning methods either require costly real-world exploration or depend on hand-crafted safety functions. Neither scales to vision-language-action models deployed in open-world physical environments. We propose SafeDojo, the first model-based safe reinforcement learning framework for vision-language-action policies designed to learn safe actions through world model-based imagination. Specifically, SafeDojo performs online reinforcement learning on top of an interactive video world model. The world model generates action-conditioned future predictions, from which a tailored ResNet success classifier estimates per-step task progress from imagined frames and a lightweight safety head predicts per-step safety costs from latent context together with the proposed action chunk, enabling simultaneous assessment of task execution and trajectory safety. The decoupled task-reward and safety-cost signals are balanced through a Lagrangian-based constrained GRPO objective, enabling coordinated improvement of task success and safety under explicit constraints. On SafeLIBERO, SafeDojo achieves the best aggregate task success, safe success, and execution efficiency among inference-time safety, model-free RL, and model-based RL baselines, with the best average safe-success rate on both levels and an 8.25 percentage-point improvement over the strongest baseline on Level I. Real-world Franka deployment further shows the best average task and safe-success rates across five tasks. Our results position world model-based safe reinforcement learning as a scalable and generalizable path toward safe embodied intelligence.

发表机构

  • State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室)
  • Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
  • Nanyang Technological University(南洋理工大学)
  • Hong Kong University of Science and Technology(香港科技大学)
  • University of Electronic Science and Technology of China(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑