arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Embodied-BenchForge:用于具身基准构建的闭环智能体工作流

Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction

Baoyang Jiang, Fengchun Zhang, Leyuan Wang, Haotian Li, Yida Wang, Zhe Ji, Jinshan Lai, Xi Ren, Danyang Li, Zheng Yang, Jianwei Hu, Qiang Ma

arXiv 2609.13082首次发表:更新:

发表机构

QiYuan Lab; University of Electronic Science and Technology of China; Beijing University of Posts and Telecommunications; Northeastern University; Beihang University; Tsinghua University(启元实验室; 电子科技大学; 北京邮电大学; 东北大学; 北京航空航天大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Embodied-BenchForge提出闭环智能体工作流,通过前向合成与后向验证修复构建具身基准,生成六个离线及一个含220任务的交互式基准,有效区分模型能力。

AI 中文摘要

智能体系统为自动化具身基准构建提供了一种有前景的方式,但现有方法通常覆盖孤立阶段,或局限于预定义的环境和任务族。更重要的是,多步骤构建会产生相互依赖的中间产物,这些产物往往未经特定于产物的验证就被传递到下游,导致局部缺陷传播到最终基准中。我们提出了Embodied-BenchForge,一个将用户指定的评估意图转化为完整具身基准产物的智能体框架。它将构建形式化为闭环基准合成,整合了前向产物合成与后向验证及修复。技能编排的产物合成将类型化且可复用的技能组合成可执行工作流,同时产物依赖图记录中间输出及其依赖关系。需求引导的验证与修复在整个构建过程中应用特定于产物的契约,并在验证失败时利用溯源触发局部重新执行或上游回滚。Embodied-BenchForge在离线EQA轨道中构建了覆盖多种具身场景的六个基准,以及在交互式具身轨道中构建了一个包含220个可执行任务的交互式基准。对代表性多模态大语言模型和具身智能体的评估表明,这些基准能够区分模型在基于观察的理解和闭环执行方面的能力。质量评估和消融实验验证了基准质量以及验证与修复的有效性,而修复和技能复用分析则展示了高效的局部恢复和跨基准可复用性。

英文摘要

Agentic systems offer a promising way to automate embodied benchmark construction, but existing approaches typically cover isolated stages or remain specialized to predefined environments and task families. More importantly, multi-step construction produces dependent intermediate artifacts that are often passed downstream without artifact-specific verification, allowing local defects to propagate into the final benchmark. We present Embodied-BenchForge, an agentic framework that transforms user-specified evaluation intents into complete embodied benchmark artifacts. It formulates construction as Closed-Loop Benchmark Synthesis, integrating forward artifact synthesis with backward verification and repair. Skill-Orchestrated Artifact Synthesis composes typed and reusable skills into executable workflows, while an artifact dependency graph records intermediate outputs and their dependencies. Requirement-Guided Verification and Repair applies artifact-specific contracts throughout construction and uses provenance to trigger local re-execution or upstream rollback when verification fails. Embodied-BenchForge constructs six benchmarks covering diverse embodied scenarios in the Offline EQA Track, together with one interactive benchmark containing 220 executable tasks in the Interactive Embodied Track. Evaluations of representative MLLMs and embodied agents show that the benchmarks distinguish model capabilities in both observation-based understanding and closed-loop execution. Quality assessment and ablations validate benchmark quality and the effectiveness of verification and repair, while repair and skill-reuse analyses demonstrate efficient localized recovery and cross-benchmark reusability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑