arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16590cs.RO

Zetta ζ:用于自进化物理智能的高效闭环具身框架

Zetta $ζ$: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence

Xin Ding, Liang Mi, Mingzhe Huang, Zixuan Wang, Chao Zhang, Zixu Hao, Fu Chen, Xiangyu Li, Yikai Zheng, Yaoyu Guo, Weijun Wang, Kun Li, Hao Wu, Yunxin Liu, Ting Cao

首次发表
浏览论文内容

中文总结 AI 辅助

Zetta ζ是一种闭环具身框架,通过时间尺度分离的循环进化评判器与恢复技能,结合Z-Infra在LIBERO-Pro、RoboCasa上实现90.8%、93.6%的成功率,为物理智能扩展提供路径。

中文摘要 AI 辅助

具身智能体越来越多地用于弥合端到端策略模型留下的差距。然而,智能体路径尚未在物理执行中实现闭环学习:现有框架在很大程度上仍为开环,在部署过程中遵循固定技能,仅在一个回合完成后才进行反思。这种事后反思无法在执行过程中进行管控,因为物理交互要求决策以超出当前大型智能体模型的频率跟踪快速变化的机器人-环境状态。我们提出Zetta,一种闭环具身框架,在冻结基础策略的同时在线进化基于代码的运行时评判器与恢复技能。通过三个时间尺度分离的循环,Zetta提供动作频率管控、部署级评判器-恢复提案,以及验证门控的技能更新。结合将智能体逻辑与异构执行资源解耦的部署基础设施Z-Infra,Zetta在当前部署预算下,在LIBERO-Pro和RoboCasa上达到了最先进的成功率,分别为90.8%和93.6%,并实现了11.1倍的推理加速;成功率随自探索经验持续扩展,学习到的技能实现零样本迁移,且出现了清晰的机器人“顿悟时刻”。这些结果表明,闭环框架自进化为可靠物理智能开辟了一条扩展路径。

英文摘要

Embodied agents are increasingly used to close the gap left by end-to-end policy models. Yet the agentic path has not realized closed-loop learning in physical execution: existing harnesses remain largely open-loop, following fixed skills during rollout and reflecting only after an episode completes. Such post-hoc reflection cannot govern execution as it unfolds, because physical interaction requires decisions to track rapidly changing robot-environment states at a frequency beyond today's large agentic models. We present Zetta, a closed-loop embodied harness that evolves code-based runtime critics and recovery skills online while keeping the base policy frozen. Through three timescale-separated loops, Zetta provides action-frequency governance, rollout-level critic-recovery proposal, and validation-gated skill updates. Together with Z-Infra, a rollout infrastructure decoupling agent logic from heterogeneous execution resources, Zetta achieves state-of-the-art success on LIBERO-Pro and RoboCasa under our current rollout budget, reaching 90.8% and 93.6%, with an 11.1x inference speedup; success continues to scale with self-exploration experience; learned skills transfer zero-shot, and clear robotic "Aha Moments" emerge. These results show that closed-loop harness self-evolution opens a scaling path for reliable physical intelligence.

发表机构

  • Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院(AIR))
  • Z-Trans AI

机构由 AI 辅助整理,请以论文原文为准。

↑