arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33807cs.ROcs.AI

CodeActionBench:评估具身操作中的智能体代码即策略

CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation

  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Yiheng Lyu, Xueying Jiang, Wenhao Li, Shijian Lu, Gongjie Zhang

AI总结:

提出CodeActionBench基准,含25个操作任务,评估通用多模态模型通过智能体代码即策略将视觉理解转化为具身操作的能力,最强配置GPT-6 Astra在675次尝试中成功率最高达73.3%,但仍有较大改进空间。

AI中文摘要:

通用多模态模型在多大程度上能够将视觉理解与推理通过可执行代码转化为具身操作?我们提出了CodeActionBench,一个包含25个操作任务的基准,通过智能体代码即策略(Code-as-Policy)来评估这一能力。在没有任务特定微调、演示、外部专家感知或抓取模块、特权场景状态或预定义任务策略的情况下,智能体应选择视觉证据,形成与任务相关的3D估计,构建操作目标,并迭代执行和修正其策略。共享的机器人API提供RGB观测、校准的几何操作、机器人反馈和有界运动,将任务相关的决策留给被评估的智能体。固定的任务实例、资源预算和隐藏的物理结果验证器支持跨模型和配置的受控比较。在九种配置和675次尝试中的广泛评估实现了从2.7%到73.3%的成功率。最强的配置,即带有Codex CLI的GPT-6 Astra,在三次尝试中至少一次解决了25个任务中的22个,展示了最佳性能,同时仍留有相当大的改进空间。轨迹分析揭示了在空间对齐、物体保持和完成判断方面的困难,包括在成功完成运动的情况下仍出现任务失败。CodeActionBench提供了一个受控的测试平台,用于衡量通用模型如何将其能力转化为操作行为,并检查该过程中的典型失败场景。

英文摘要:

How well can general-purpose multimodal models turn visual understanding and reasoning into embodied manipulation via executable code? We introduce CodeActionBench, a benchmark of 25 manipulation tasks that evaluates this capability through agentic Code-as-Policy. Without task-specific fine-tuning, demonstrations, external specialist perception or grasp modules, privileged scene state, or predefined task policies, agents should select visual evidence, form task-relevant 3D estimates, construct manipulation targets, and iteratively execute and revise their policies. A shared robot API provides RGB observations, calibrated geometric operations, robot feedback, and bounded motion, leaving task-dependent decisions to the evaluated agent. Fixed task instances, resource budgets, and a hidden physical-outcome verifier support controlled comparisons across models and harness configurations. Extensive evaluations across nine configurations and 675 attempts achieve success rates ranging from 2.7% to 73.3%. The strongest configuration, GPT-6 Astra with Codex CLI, solves 22 of 25 tasks at least once in three attempts, demonstrating the best performance while still leaving substantial room for improvement. Trajectory analyses reveal difficulties in spatial alignment, object retention, and completion judgment, including task failures despite successfully completed motions. CodeActionBench provides a controlled testbed for measuring how general-purpose models translate their capabilities into manipulation behavior and for examining typical failure scenarios in that process.

补充信息

↑