arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhysClaw-0:一种通过语言修正实现机器人自主的共生智能体系统

Zero2Skill: Bootstrapping Robot Skills through Autonomous Data Collection, Training, and Deployment

Boyuan Wang, Zhenyuan Zhang, Zhiqin Yang, Peijun Gu, Shuya Wang, Xiaofeng Wang, Xianghui Ze, Yifan Chang, Guosheng Zhao, Jiangnan Shao, Guan Huang, Hengyu Liu, Yonggang Zhang, Wei Xue, Chunyuan Guan, Chenglin Pu, Yike Guo, Xingang Wang, Zheng Zhu

arXiv 2607.14047首次发表:更新:

发表机构

GigaAI; University of Chinese Academy of Sciences; Hong Kong University of Science and Technology; University of Leeds; Cornell University; Tsinghua University; Nanjing University of Science and Technology; The Chinese University of Hong Kong; FAWTD(极佳科技; 中国科学院大学; 香港科技大学; 利兹大学; 康奈尔大学; 清华大学; 南京理工大学; 香港中文大学; 一汽技术开发部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对自主数据收集问题,提出PhysClaw-0共生智能体系统,通过跨轮保留和重用修正、自主收集验证等方式,在真实机器人测试中减少人力时间,提高成功率,并提升验证者与人类一致性。

AI 中文摘要

自主数据收集决定了用于操作策略学习的真实世界轨迹的数量和质量。现有流程通过自我重置、VLM验证或语言引导修正来减少人力,但相同故障复发时需重新进行情节范围内的修复,监督成本随会话长度而非不同问题数量增长。我们提出PhysClaw-0,这是一种人机共生智能体系统,修正可跨轮保留和重用。收集循环自主收集、验证和重置,仅在阶段耗尽明确重试预算时暂停以等待远程操作员。LLM解析器将自然语言话语映射到存储在纠正记忆中的结构化调整,因此相同条件下已解决的故障模式通常无需再次修正。在真实机器人桌面清理测试平台上,PhysClaw-0在将人类工作时间减少到16%的同时,达到了遥操作情节成功率。语言修正提高了所有四种评估设置下验证者与人类的一致性,并将平均单次尝试成功率从12.5%提高到47.5%(手臂选择方面从20.)

英文摘要

Autonomous data collection governs the volume and quality of real-world trajectories for manipulation policy learning. Existing pipelines reduce human effort via self-resetting, VLM verification, or language-guided correction, yet episode-scoped fixes must be reissued whenever the same failure recurs, so oversight cost grows with session length rather than with the number of distinct problems. We present Zero2Skill, a human-robot symbiotic agentic system in which corrections are retained and reused across rounds. The collection loop collects, verifies, and resets autonomously, pausing for a remote operator only when a phase exhausts an explicit retry budget. An LLM parser maps each natural-language utterance to a structured adjustment stored in Corrective Memory, so addressed failure modes typically need not be corrected again under the same conditions. On a real-robot desktop-clearing testbed, Zero2Skill matches teleoperation episode success while reducing human working time to 16%. Language corrections improve verifier-human agreement in all four evaluated settings and raise average single-attempt success from 12.5% to 47.5% (arm-selection: 20.0% to 50.0%). Policies fine-tuned on Zero2Skill data match teleoperation-trained policy success at a fraction of collection human cost.

CommentsWebPage: https://open-gigaai.github.io/Zero2Skill

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑