arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI科学家能否在运行时进行协调?

Can AI Scientists Coordinate at Runtime?

Zijian Liu, Yangzhixin Luo, Junyu Lu, Yi Li, Yu Chen, David Xu, William F. Shen, Xinchi Qiu, Xisen Wang

arXiv 2610.00980首次发表:更新:

发表机构

University of Oxford; King Abdullah University of Science and Technology; University of Sydney; University of Cambridge(牛津大学; 阿卜杜拉国王科技大学; 悉尼大学; 剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨AI科学家能否在运行时协调,提出运行时智能体协调(RAC)方法,在ResearchClawBench上评估显示运行时选择提升性能,但额外协调机制在预算受限时效果有限。

AI 中文摘要

多智能体AI科学家在多种任务中展现出不断提升的性能。然而,一种常见的方法是设计时的智能体编排,这通常依赖于固定的工作流程。相比之下,人类科学家会在运行时协调并调整他们的分工。因此,我们提出疑问:AI科学家是否也能在运行时进行协调?为此,我们引入了运行时智能体协调(Runtime Agent Coordination, RAC),该方法在执行过程中从现有的AI科学家宿主中选择智能体,分配范围明确的工作合同,并提供基于产物的验证。验证会为后续智能体提供信息,而不会阻塞转换或丢弃产物。我们在ResearchClawBench上对Agent Laboratory、EvoScientist和ARK进行了单种子探索性评估,在宿主校准的预算下保留了宿主的模型、工具和权限。四个累积条件分别区分了原生执行、运行时通信、运行时选择以及合同与验证的组合添加。运行时选择为每个宿主带来了观察到的最高平均分数;添加合同和验证会降低这些平均值,且相对于原生执行的结果取决于宿主。这些结果推动了运行时协调的发展,同时也揭示了在预算受限的情况下,额外协调机制的局限性。代码可在以下网址获取:https://this-url。

英文摘要

Multi-agent AI scientists have shown improving performance across a diverse range of tasks. Yet a common approach is design-time agentic orchestration, which typically relies on fixed workflows. In contrast, human scientists coordinate and adjust their division of labor at runtime. We therefore ask: can AI scientists also coordinate at runtime? To this end, we introduce Runtime Agent Coordination (RAC), which selects agents from existing AI-scientist hosts during execution, assigns scoped work contracts, and provides artifact-grounded verification. Verification informs subsequent agents without blocking transitions or discarding artifacts. We conduct a single-seed exploratory evaluation across Agent Laboratory, EvoScientist, and ARK on ResearchClawBench, preserving host models, tools, and permissions under host-calibrated budgets. Four cumulative conditions separate native execution, runtime communication, runtime selection, and the combined addition of contracts and verification. Runtime selection yields the highest observed mean score for each host; adding contracts and verification reduces these means, with host-dependent outcomes relative to native execution. These results motivate runtime coordination while exposing the limits of additional coordination mechanisms under constrained budgets. Code is available at https://github.com/systemind-team/Runtime-AI-Scientist.

Comments35 pages (9 pages main text), 4 figures, 10 tables. Code: https://github.com/systemind-team/Runtime-AI-Scientist

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑