发表机构
Galbot Team(Galbot 团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统评估GPT-6 Astra在六个机器人领域的具身策略能力,发现其在任务决策上表现优异,但物理控制可靠性不足,且推理延迟高,制约实际应用。
AI 中文摘要
GPT-6 Astra展现出生成数值机器人动作的卓越能力,将其角色扩展至高层规划之外。为评估Astra作为通用具身策略的能力,我们在六个领域进行了全面评估,考察直接控制、与学习策略的协作以及反馈驱动的适应。在夹爪操作中,Astra能够纠正任务目标并为后续策略执行准备接触条件;与π0.5的混合控制在评估的RoboDojo子集上达到48%的成功率。在灵巧操作中,混合控制在十次经验引导的DexJoCo试验中达到50%的成功率,而直接手内控制则难以协调手指接触。在移动操作中,混合控制在评估的RoboCasa365上达到38.7%的成功率。在导航中,Astra在我们的局部比较中领先,在RxR指令跟随上达到92%的成功率,在HM3D物体搜索上达到82%的成功率,尽管搜索产生了大量绕路。在运动生成中,密集运动参考生成仍不可靠:在单个障碍物课程上的五次连续尝试中,没有一次到达目标,尽管在稳定性和前进进度上有所改善。在人形移动操作中,Astra在30个HumanoidBench任务中的13个上超过了使用预训练全身控制器的基线方法。这些发现揭示了有用的任务决策与可靠的物理控制之间的差距。推理延迟进一步限制了实际控制:在每种条件下的50个RoboDojo实例中,策略辅助和直接控制分别消耗6.248亿和11.32亿个令牌。一次30秒的运动运行需要250次模型调用,平均每次39.86秒,推理期间物理暂停。
英文摘要
GPT-6 Astra exhibits a remarkable ability to generate numerical robot actions, extending its role beyond high-level planning. To assess Astra's capabilities as general-purpose embodied policies, we conduct comprehensive evaluations across six domains, examining direct control, cooperation with learned policies, and feedback-driven adaptation. In gripper manipulation, Astra can correct task targets and prepare contact conditions for subsequent policy execution; hybrid control with π0.5 achieves 48% success on the evaluated RoboDojo subset. In dexterous manipulation, hybrid control achieves 50% success in ten experience-guided DexJoCo trials, while direct in-hand control struggles to coordinate finger contacts. In mobile manipulation, hybrid control reaches 38.7% success on the evaluated RoboCasa365. In navigation, Astra leads our local comparisons, reaching 92% success on RxR instruction following and 82% on HM3D object search, although search incurs substantial detours. In locomotion, dense motion-reference generation remains unreliable: none of five sequential attempts on a single obstacle course reaches the goal, despite improvements in stability and forward progress. In humanoid loco-manipulation, Astra exceeds baseline methods on 13 of 30 HumanoidBench tasks with pretrained whole-body controllers. These findings reveal a gap between useful task decisions and reliable physical control. Inference latency further constrains practical control: across 50 RoboDojo instances per condition, policy-assisted and direct control consume 624.8 million and 1.132 billion tokens. A 30-second locomotion run requires 250 model calls averaging 39.86 seconds each, with physics paused during inference.