arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39507cs.RO

LIBERO-Agent:评估通用智能体在直接具身操作中的表现

LIBERO-Agent: Evaluating General-Purpose Agents for Direct Embodied Manipulation

  • Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身智能研究院)
  • Shanghai Innovation Institute(上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

Zijie Diao, Yitong Chen, Sicheng Xie, Tianyi Lu, Wujian Peng, Guojin Zhong, Houze Xu, Ziyi Ye, Zuxuan Wu, Yu-Gang Jiang

中文总结 AI 辅助

该论文提出LIBERO-Agent基准,用于评估通用智能体在具身操作中的能力,发现其存在可靠性差距,且GPT-6 Astra表现最佳,主要优势在机制交互,失败源于跨阶段干扰和几何误差。

中文摘要 AI 辅助

通用智能体能够规划、使用工具并根据反馈修正自身行为,但这些能力能否从数字环境迁移到具身操作中仍不清楚。为探究这一问题,我们引入了LIBERO-Agent,一个用于在机器人操作任务中评估这些智能体的智能体原生基准。LIBERO-Agent并非要求智能体提交任务级Python控制程序或通过高级机器人技能进行操作,而是提供一个交互式机器人环境,其中智能体可以选择检查哪些观测、使用自身工具处理这些观测,并发出原生动作命令。LIBERO-Agent将200个任务整合到一个通用交互框架中,并提供了一个包含30个任务的主要套件,该套件将感知、短时程执行和长时程组合分开。结果揭示了显著的可可靠性差距:尽管智能体在感知和简单短时程任务上表现良好,但在困难短时程和长时程任务上性能大幅下降。更丰富的观测改善了短时程操作,而示范的好处取决于智能体和格式。在这些智能体中,GPT-6 Astra取得了最强的整体性能。进一步分析显示,其主要优势在于机制交互,尤其是在需要持续物理接触时,而其剩余失败源于跨阶段干扰和几何误差。

英文摘要

General-purpose agents can plan, use tools, and revise their behavior from feedback, but it remains unclear whether these capabilities transfer from digital environments to embodied manipulation. To investigate this question, we introduce LIBERO-Agent, an agent-native benchmark for evaluating these agents in robot manipulation tasks. Rather than asking agents to submit task-level Python control programs or operate through high-level robot skills, LIBERO-Agent provides an interactive robotic environment where agents can select which observations to inspect, process them with their own tools, and issue native action commands. LIBERO-Agent integrates 200 tasks into a common interaction framework and provides a 30-task primary suite that separates perception, short-horizon execution, and long-horizon composition. Results reveal a pronounced reliability gap: while agents perform well on perception and easy short-horizon tasks, their performance degrades substantially on hard short-horizon and long-horizon tasks. Richer observations improve short-horizon manipulation, while demonstration benefits depend on the agent and format. Among these agents, GPT-6 Astra achieves the strongest overall performance. Further analysis shows its major advantage lies in mechanism interaction, especially when sustained physical contact is needed, while its remaining failures stem from cross-stage interference and geometric errors.

↑