arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24170cs.CV

一个出人意料的机器人策略:GPT-6 Astra 在 RoboDojo 及更广泛任务上的早期评估

An Unexpected Robot Policy: Early Evaluations of GPT-6 Astra on RoboDojo and Beyond

Wenbo Zhang, Kaixuan Wang, Yutao Ouyang, Xiaoyu Huang, Liyang Li, Kailun Su, Weiyang Jin, Wenhao Chai, Haotian Liang, Zhiyang Dou, Yue Chen, Tianxing Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估三个LLM在42个RoboDojo任务上的操作能力,发现GPT-6 Astra以22.48%成功率领先,但精度和动态控制仍是局限。

中文摘要 AI 辅助

具身人工智能系统通常被组织为系统1和系统2。系统1通常是一个预训练的策略,以高频率生成动作,而系统2通常被实例化为一个支持视觉的语言模型,用于高层规划。我们探究一个大型语言模型(LLM)是否可以在没有任务特定微调的情况下,充当机器人操作策略。我们将这种设置称为“LLM作为策略”。我们在所有42个RoboDojo任务上评估了三个LLM,并将它们的得分与40个公开策略进行比较。Astra和GPT-5.5使用官方的每个任务50个回合的协议;DeepSeek-Flash使用每个任务10个回合。GPT-6 Astra在2100次试验中实现了22.48%的平均成功率和28.97的得分,排名高于所有公开条目。然而,GPT-5.5和DeepSeek-Flash在相同的后处理下仅达到0.88%和1.92%的平均成功率。我们发现Astra表现出一种极度两极化的能力轮廓。它很好地泛化到需要语义理解但不需要高精度控制的任务。相反,它在需要精度、动态控制或复杂双臂协调的任务上表现不佳。上下文内实验显示,一次性演示没有带来总体收益,而选定的交互轨迹显示在扰动下回合内修正。总体而言,所评估的LLM在操作性能上差异显著。Astra脱颖而出,为通用操作模型的潜力提供了初步证据,尽管在评估设置中,可靠的精度和动态控制仍然是局限性。

英文摘要

Embodied AI systems are often organized into System 1 and System 2. System 1 is typically a pretrained policy that generates actions at high frequency, whereas System 2 is often instantiated as a vision-enabled language model for high-level planning. We ask whether a large language model (LLM) can act as the policy for robot manipulation without task-specific finetuning. We call this setting LLM as policy. We evaluate three LLMs on all 42 RoboDojo tasks and compare their scores with 40 public policies. Astra and GPT-5.5 use the official 50-episode-per-task protocol; DeepSeek-Flash uses 10 episodes per task. GPT-6 Astra achieves 22.48% average success rate and 28.97 Score over 2,100 trials, ranking above every public entry. Yet GPT-5.5 and DeepSeek-Flash reach only 0.88% and 1.92% average success rate with the same post-processing. We find that Astra exhibits a sharply polarized capability profile. It generalizes well to tasks that require semantic understanding but not high-precision control. In contrast, it performs poorly on tasks that require precision, dynamic control, or complex bimanual coordination. In-context experiments show no aggregate benefit from one-shot demonstrations, while selected interaction traces show within-episode corrections under perturbations. Overall, the evaluated LLMs vary substantially in manipulation performance. Astra stands out and provides initial evidence for the potential of a general-purpose manipulation model, although reliable precision and dynamic control remain limitations in the evaluated setting.

发表机构

  • RoboProbe
  • RoboDojo
  • The University of Hong Kong(香港大学)
  • Tsinghua University(清华大学)
  • University of California, Berkeley(加州大学伯克利分校)
  • Princeton University(普林斯顿大学)
  • Massachusetts Institute of Technology(麻省理工学院)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑