发表机构
Harvard University; Georgia Institute of Technology(哈佛大学; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RLE-Bench提出一个覆盖四种机器人开发工作流的基准,通过多指标评估编码智能体的工程能力,并引入RLE指数以系统比较其多维性能。
AI 中文摘要
编码智能体正开始超越纯数字任务,涉足物理世界的挑战,尤其是在机器人领域。然而,现有的机器人基准主要关注单个工件(如策略或控制器)的性能,对编码智能体更广泛的工程能力的覆盖有限。现实世界的机器人技术不仅限于控制:智能体必须在资源约束下构建、集成、诊断和改进异构工件,并从多模态反馈中进行推理。为了评估这些更广泛的能力,我们引入了RLE-Bench,一个涵盖四种代表性机器人开发工作流的机器人学习任务基准:交互控制、策略学习、感知与估计,以及机械设计。我们使用多样化的任务特定指标来评估编码智能体提交的工件,从智能体实现的成功率到其训练的策略、构建的测试平台以及设计的机械结构。我们将这些指标聚合为一个总体RLE指数,并报告工作流特定的能力概况,从而能够在多个能力维度上系统比较编码智能体的能力。除了性能排名外,我们还进行了深入的案例研究,检查智能体在代表性任务上的行为,突出当前的能力和局限性,并指出机器人任务为未来智能体训练提供的机会。
英文摘要
Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited coverage of coding agents' broader engineering capabilities. Real-world robotics extends beyond control: agents must build, integrate, diagnose, and improve heterogeneous artifacts under resource constraints and reason from multimodal feedback. To evaluate these broader capabilities, we introduce RLE-Bench, a benchmark of robot-learning tasks spanning four representative robotics development workflows: interactive control, policy learning, perception and estimation, and mechanical design. We use diverse task-specific metrics to evaluate the artifacts submitted by the coding agents, from the success rate the agents achieved to the policy agents trained, the harness agent built, and the mechanical structures the agent designed. We aggregate these metrics into an overall RLE Index and report workflow-specific capability profiles, enabling systematic comparison of coding agents' capabilities across multiple capability dimensions. Beyond performance ranks, we also conduct in-depth case studies examining agent behavior on representative tasks, highlighting both current capabilities and limitations, and pointing to the opportunities robotics tasks have to offer for future agent training.
Comments28 pages, 17 figures. Project website: https://rle-bench.github.io/