AI 中文总结
本研究利用GPT-6-Astra控制机器人完成电梯按钮任务,通过提供身体知识、经验复用和技能涌现,显著提升效率,并验证了模拟到真实的迁移能力。
AI 中文摘要
通用多模态智能体能够编写机器人控制程序,但反复探索和模型介导的动作选择可能使执行变得缓慢。我们研究了外部身体知识、成功经验和可执行技能如何改善由GPT-6-Astra控制的XLeRobot在模拟和物理电梯按钮任务中的表现。在30次固定起始位置的模拟试验中,完整的机器人几何和相机信息相比仅具有通用控制接口且无先前经验的基线,平均完成时间减少了57.4%;带有同步动作和状态记录的图像在无需额外身体资产的情况下减少了68.6%。在起始位置偏移10-100厘米的九对比较(18次试验)中,在原始起始位置记录的经验相比无经验平均时间减少了58-63%,展示了对测试的新起始位置的泛化能力。在经验实验期间,GPT-6-Astra自发生成了一个简短的视觉反馈程序。研究者重构的版本在27次模拟试验中平均局部任务时间减少了29-31%。最后,使用操作员确认的按钮接触的12次真实机器人试验展示了sim2real复用:在共享标称起始位置,模拟XML资产和模拟经验分别使平均时间减少了53.0%和49.9%;真实经验也迁移到了两个新的起始位置。这些结果表明了一种围绕GPT-6-Astra构建通用操控实验的实用方法:提供机器可读的身体描述和同步演示,将智能体生成的有用反馈例程转化为可复用技能,同时智能体根据当前图像调整动作。我们在该https URL发布了所有任务提示、试验级实验数据和获得的技能实现。
英文摘要
General-purpose multimodal agents can write robot-control programs, but repeated exploration and model-mediated action selection can make execution slow. We study how external body knowledge, successful experience, and executable skills improve an XLeRobot controlled by GPT-6-Astra in a simulated and a physical elevator-button task. In 30 fixed-start simulation trials, complete robot geometry and camera information reduce mean completion time by 57.4% relative to a baseline with only the common control interface and no prior experience; images with synchronized action and state records reduce it by 68.6% without additional body assets. In nine paired comparisons (18 trials) at starts displaced by 10-100 cm, experience recorded at the original start reduces mean time by 58-63% relative to no experience, demonstrating generalization to the tested new starting positions. During experience experiments, GPT-6-Astra spontaneously generates a short visual-feedback program. Researcher-refactored versions reduce mean local-task time by 29-31% in 27 simulation trials. Finally, 12 real-robot trials using operator-confirmed button contact demonstrate sim2real reuse: at a shared nominal start, simulation XML assets and simulation experience reduce mean time by 53.0% and 49.9%, respectively; real experience also transfers to two new starts. These results suggest a practical way to build general-purpose manipulation experiments around GPT-6-Astra: supply machine-readable body descriptions and synchronized demonstrations, and turn useful agent-generated feedback routines into reusable skills, while the agent adapts actions from current images. We release all task prompts, trial-level experimental data, and acquired skill implementations at https://github.com/hesd10/astra-robot-sim2real.
Comments23 pages, 10 figures, 6 tables. Code, data, prompts, and skills: https://github.com/hesd10/astra-robot-sim2real