一个出人意料的机器人策略:GPT-6 Astra 在 RoboDojo 及更广泛任务上的早期评估
An Unexpected Robot Policy: Early Evaluations of GPT-6 Astra on RoboDojo and Beyond
浏览论文内容
中文总结 AI 辅助
本研究评估三个LLM在42个RoboDojo任务上的操作能力,发现GPT-6 Astra以22.48%成功率领先,但精度和动态控制仍是局限。
中文摘要 AI 辅助
具身人工智能系统通常被组织为系统1和系统2。系统1通常是一个预训练的策略,以高频率生成动作,而系统2通常被实例化为一个支持视觉的语言模型,用于高层规划。我们探究一个大型语言模型(LLM)是否可以在没有任务特定微调的情况下,充当机器人操作策略。我们将这种设置称为“LLM作为策略”。我们在所有42个RoboDojo任务上评估了三个LLM,并将它们的得分与40个公开策略进行比较。Astra和GPT-5.5使用官方的每个任务50个回合的协议;DeepSeek-Flash使用每个任务10个回合。GPT-6 Astra在2100次试验中实现了22.48%的平均成功率和28.97的得分,排名高于所有公开条目。然而,GPT-5.5和DeepSeek-Flash在相同的后处理下仅达到0.88%和1.92%的平均成功率。我们发现Astra表现出一种极度两极化的能力轮廓。它很好地泛化到需要语义理解但不需要高精度控制的任务。相反,它在需要精度、动态控制或复杂双臂协调的任务上表现不佳。上下文内实验显示,一次性演示没有带来总体收益,而选定的交互轨迹显示在扰动下回合内修正。总体而言,所评估的LLM在操作性能上差异显著。Astra脱颖而出,为通用操作模型的潜力提供了初步证据,尽管在评估设置中,可靠的精度和动态控制仍然是局限性。
英文摘要
Embodied AI systems are often organized into System 1 and System 2. System 1 is typically a pretrained policy that generates actions at high frequency, whereas System 2 is often instantiated as a vision-enabled language model for high-level planning. We ask whether a large language model (LLM) can act as the policy for robot manipulation without task-specific finetuning. We call this setting LLM as policy. We evaluate three LLMs on all 42 RoboDojo tasks and compare their scores with 40 public policies. Astra and GPT-5.5 use the official 50-episode-per-task protocol; DeepSeek-Flash uses 10 episodes per task. GPT-6 Astra achieves 22.48% average success rate and 28.97 Score over 2,100 trials, ranking above every public entry. Yet GPT-5.5 and DeepSeek-Flash reach only 0.88% and 1.92% average success rate with the same post-processing. We find that Astra exhibits a sharply polarized capability profile. It generalizes well to tasks that require semantic understanding but not high-precision control. In contrast, it performs poorly on tasks that require precision, dynamic control, or complex bimanual coordination. In-context experiments show no aggregate benefit from one-shot demonstrations, while selected interaction traces show within-episode corrections under perturbations. Overall, the evaluated LLMs vary substantially in manipulation performance. Astra stands out and provides initial evidence for the potential of a general-purpose manipulation model, although reliable precision and dynamic control remain limitations in the evaluated setting.
发表机构
- RoboProbe
- RoboDojo
- The University of Hong Kong(香港大学)
- Tsinghua University(清华大学)
- University of California, Berkeley(加州大学伯克利分校)
- Princeton University(普林斯顿大学)
- Massachusetts Institute of Technology(麻省理工学院)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。