arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27308cs.RO

EmbodiedSWE:面向长时程灵巧机器人的编码智能体

EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics

Zeyu Shen, Haoxiang You, Yilang Liu, Zhicheng Zheng, Lihan Zha, Kashu Yamazaki, Mingtong Zhang, Suning Huang, Jiankai Sun, Qianzhong Chen, Lucy He, Kaiyuan Liu,… 展开作者

Zeyu Shen, Haoxiang You, Yilang Liu, Zhicheng Zheng, Lihan Zha, Kashu Yamazaki, Mingtong Zhang, Suning Huang, Jiankai Sun, Qianzhong Chen, Lucy He, Kaiyuan Liu, Haoran Chang, Katerina Fragkiadaki, Dhruv Shah, Mac Schwager, Peter Henderson, Ian Abraham, Canwen Xu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出EmbodiedSWE框架,利用编码智能体解决长时程灵巧机器人任务,并通过生成多样化仿真轨迹训练VLA,实现从仿真到真实机器人的策略泛化。

中文摘要 AI 辅助

我们研究了面向长时程、灵巧机器人的编码智能体,并探究其解决方案能否为学习通用机器人策略提供可扩展的监督。为验证这一点,我们开发了EMBODIEDSWE-BENCH,这是一个面向编码智能体的仿真基准,涵盖接触丰富的操作、可变形物体以及需要长达半小时连续交互的长时程任务。我们发现,前沿编码智能体能够解决复杂的长时程任务,并将先前解决方案跨任务和跨具身迁移。我们还设计了辅助工具,帮助智能体更有效地解决这些任务。然而,所得解决方案需要大量的迭代交互,且通常针对单个任务实例特化。因此,我们引入EMBODIEDSWE-GEN,它将编码智能体的单一解决方案扩展为用于训练视觉-语言-动作模型(VLA)的大规模多样化轨迹。VLA性能随生成的示范数量增加而提升,且智能体辅助的多样化增强了在保留任务变体上的泛化能力。我们还展示,仅使用编码智能体生成的仿真示范进行微调的VLA,能在真实机器人上完成长时程任务。综合而言,我们的框架利用编码智能体解决复杂机器人任务,并将验证过的解决方案转化为机器人策略的可扩展监督。

英文摘要

We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervision for learning general robot policies. To test this, we develop EMBODIEDSWE-BENCH, a simulation benchmark for coding agents spanning contact-rich manipulation, deformable objects, and long-horizon tasks requiring up to half an hour of continuous interaction. We find that frontier coding agents can solve complex long-horizon tasks and transfer prior solutions across both tasks and embodiments. We also design supporting tools that help agents more effectively solve these tasks. However, the resulting solutions require substantial iterative interaction and are typically specialized to individual task instances. We therefore introduce EMBODIEDSWE-GEN, which expands a single solution from coding agent into large diverse trajectories for training a VLA. VLA performance improves with more generated demonstrations, and agent-aided diversification improves generalization to held-out task variations. We also show that a VLA finetuned solely on coding-agent-generated simulation demonstrations completes a long-horizon task on real robot. Together, our framework uses coding agents to solve complex robotics tasks and turn verified solutions into scalable supervision for robot policies.

发表机构

  • ByteDance Seed(字节跳动)
  • Yale University(耶鲁大学)
  • Princeton University(普林斯顿大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • Stanford University(斯坦福大学)
  • University of California, Los Angeles(加州大学洛杉矶分校)
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

相关深度报道

↑