发表机构
University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究在智能体机器人技术背景下,借助带Fable 5的Claude Code编码智能体,无需演示数据解决Push-T等机器人操作任务,其策略成功率达100%,步骤比基准扩散策略少46%,还支持多字母操作及跨实体机械臂仿真。
AI 中文摘要
Push-T是从人类演示中学习操作策略的标志性基准任务,机器人必须使用单个接触点将T形块推至目标姿态。在这篇短文里,我们在新兴的智能体机器人技术背景下重新审视Push-T任务,提示一个大型语言模型(LLM)编码智能体——带有Fable 5的Claude Code——创建无需任何演示数据的算法解决方案。我们研究智能体编码循环解决Push-T任务的有效性,并将得到的代码作为策略与视觉运动模仿学习策略进行比较。结果表明,该智能体找到了2D gym仿真环境,利用仿真实验学习推动力学,迭代优化后达到100%成功率,所需步骤比用200个人类演示训练的最佳扩散策略少46%。该编码智能体还通过自生成课程解决了从T到完整字母表(Push-A至Push-Z)的扩展任务,并在具有视觉反馈的跨实体3D仿真中生成了用于Franka和UR5机械臂的仿真代码。视频、策略和详细信息将发布在网上。
英文摘要
Push-T is an iconic benchmark for learning manipulation policies from human demonstrations. The robot must use a single point of contact to push a T-shaped block into a target pose. In this short paper, we revisit the Push-T task in the context of emerging advances in Agentic Robotics where an LLM coding agent -- Claude Code with Fable 5 -- is prompted to create an algorithmic solution that does not require any demonstration data. We study how effective the agentic coding loop can solve the Push-T task, and compare the resulting code as policy with the visuomotor imitation learning policy. Results suggest that the agent found the 2D gym simulation online, and used sim experiments to learn push mechanics, iteratively optimizing to achieve 100% success rate using 46% fewer steps than the best diffusion policy trained with 200 human demonstrations. The coding agent also solve extensions from T to the full alphabet (Push-A to Push-Z) using a self generated curriculum and generated simulation code for the Franka and UR5 robot arms in 3D cross-embodiment simulations with visual feedback. Videos, policies and details will be posted online.