arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型(LLM)能否设计高质量实验?一项针对自主实验设计的全面系统基准测试

Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

Zejun Liu, Jian Wu, Ru Peng, Yuliang Ji, Dongyuan Li, Renhe Jiang, Yue Zhang

arXiv 2608.03501首次发表:更新:

AI 中文总结

本研究构建SCOPE基准评估LLM的自主实验设计能力,发现多数LLM无法直接设计高质量实验,低层配置存在瓶颈,进而提出OptED工作流缓解该瓶颈。

AI 中文摘要

面向研究的人工智能(AI4Research)利用人工智能实现科学工作流的自动化与优化。实验设计是研究过程的关键阶段,但现有研究主要聚焦代码实现与执行,忽视了该阶段的重要性,且不存在评估AI开展系统性实验设计能力的基准。为填补这一空白,我们提出SCOPE,即科学综合规划评估基准,该基准由顶会(如ICML、NeurIPS、ICLR)19个研究领域的300篇高质量最新论文构建,从两个维度评估大语言模型(LLM):一是高层规划完整性(主实验、消融实验与分析实验),二是低层配置准确性与合理性(数据集、基准与指标)。基准测试揭示三项发现:(1)多数LLM无法直接设计高质量实验;(2)所有LLM在低层配置上均存在性能瓶颈;(3)搜索模式无法提升设计质量。此外,为应对这些挑战,我们提出OptED,一种优化基于LLM的实验设计的新型智能体工作流,该工作流通过阶段隔离、工具增强与基于规则的约束增强基于LLM的实验规划,有效缓解了配置瓶颈。

英文摘要

AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focused primarily on code implementation and execution, overlooking the importance of this stage, and no benchmark exists to evaluate AI's ability to conduct systematic experiment design. To bridge this gap, we propose SCOPE, a Scientific COmprehensive Planning Evaluation Benchmark constructed from 300 high-quality latest papers across 19 research domains from top-tier venues (e.g., ICML, NeurIPS, and ICLR),evaluating LLMs on two dimensions: High-Level planning completeness (main, ablation, and analysis experiments) and Low-Level configuration accuracy and rationality (datasets, baselines, and metrics). Benchmarking reveals three findings: (1) most LLMs cannot directly design high-quality experiments; (2) all LLMs exhibit a performance bottleneck in low-level configuration; and (3) search mode does not improve design quality. Furthermore, to address these challenges, we propose OptED, a novel agentic workflow to optimize LLM-based experimental design, that enhances LLM-based experimental planning through stage isolation, tool augmentation, and rule-based constraints, effectively alleviating the configuration bottleneck.

Comments32 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑