arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SciHorizon-eLab:面向科学具身智能体可扩展基准测试的智能体式协议到任务编译器

SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents

Maokai Qin, Chuan Qin, Qi Zhang, Dianyu Liu, Zirui Liu, Hongting Niu, Yuanchun Zhou, Hengshu Zhu

arXiv 2609.30971首次发表:更新:

发表机构

Data Intelligence for Scientific Innovation Lab, Computer Network Information Center, Chinese Academy of Sciences; Beihang University(中国科学院计算机网络信息中心科学智能创新数据实验室; 北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对科学具身智能体缺乏可扩展评估环境的问题,提出SciHorizon-eLab编译器,将自然语言协议自动编译为可执行具身任务,构建含300个认证任务的基准,最强策略成功率仅49.7%,揭示人机协调弱点。

AI 中文摘要

具身智能体为实现科学实验自动化提供了一条有前景的途径,但其进展受到缺乏可靠且系统的评估环境的制约。现有的基于模拟的实验室基准在很大程度上依赖于手动任务工程,这使得难以在规模上系统地将多样化的科学协议编译为可执行且可验证的具身任务。为应对这一挑战,我们引入了SciHorizon-eLab,一种智能体式协议到任务编译器,它将科学具身任务构建表述为一个编译问题。给定科学实验的自然语言协议,SciHorizon-eLab通过语义接地、可执行任务合成和多阶段基于模拟的认证,逐步将实验室协议编译为保持语义的具身任务。该系统生成语义接地的环境、可执行的操作程序和步骤级成功规范,同时支持专家演示和执行轨迹的可复现生成。利用这一流水线,我们进一步构建了\BenchName,一个包含300个涵盖多样化实验室操作场景的已认证任务的即用型基准。它支持HIL任务执行、可复现的专家演示生成和有序的步骤级评估。在代表性任务中,最强的策略平均成功率仅为49.7%,进一步的评估揭示了人类与具身智能体协调方面的显著弱点。我们在https URL上公开发布了代码、基准数据和评估工具包。

英文摘要

Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile diverse scientific protocols into executable and verifiable embodied tasks at scale. To address this challenge, we introduce SciHorizon-eLab, an agentic protocol-to-task compiler that formulates scientific embodied task construction as a compilation problem. Given a natural-language protocol of scientific experiments, SciHorizon-eLab progressively compiles laboratory protocols into semantic-preserving embodied tasks through semantic grounding, executable task synthesis, and multi-stage simulation-based certification. The system generates semantically grounded environments, executable manipulation programs, and step-level success specifications, while enabling reproducible generation of expert demonstrations and execution traces. Using this pipeline, we further construct \BenchName, a ready-to-use benchmark comprising 300 certified tasks across diverse laboratory operations. It supports HIL task execution, reproducible expert-demonstration generation, and ordered step-level evaluation. Across representative tasks, the strongest policy attains an average success rate of only 49.7%, with further evaluations revealing pronounced weaknesses in human and embodied agent coordination. We publicly release the code, benchmark data, and evaluation toolkit at https://github.com/SciHorizon-elab/SciHorizon-elab.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑