AI 中文总结
本研究构建SCOPE基准评估LLM的自主实验设计能力,发现多数LLM无法直接设计高质量实验,低层配置存在瓶颈,进而提出OptED工作流缓解该瓶颈。
AI 中文摘要
面向研究的人工智能(AI4Research)利用人工智能实现科学工作流的自动化与优化。实验设计是研究过程的关键阶段,但现有研究主要聚焦代码实现与执行,忽视了该阶段的重要性,且不存在评估AI开展系统性实验设计能力的基准。为填补这一空白,我们提出SCOPE,即科学综合规划评估基准,该基准由顶会(如ICML、NeurIPS、ICLR)19个研究领域的300篇高质量最新论文构建,从两个维度评估大语言模型(LLM):一是高层规划完整性(主实验、消融实验与分析实验),二是低层配置准确性与合理性(数据集、基准与指标)。基准测试揭示三项发现:(1)多数LLM无法直接设计高质量实验;(2)所有LLM在低层配置上均存在性能瓶颈;(3)搜索模式无法提升设计质量。此外,为应对这些挑战,我们提出OptED,一种优化基于LLM的实验设计的新型智能体工作流,该工作流通过阶段隔离、工具增强与基于规则的约束增强基于LLM的实验规划,有效缓解了配置瓶颈。
英文摘要
AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focused primarily on code implementation and execution, overlooking the importance of this stage, and no benchmark exists to evaluate AI's ability to conduct systematic experiment design. To bridge this gap, we propose SCOPE, a Scientific COmprehensive Planning Evaluation Benchmark constructed from 300 high-quality latest papers across 19 research domains from top-tier venues (e.g., ICML, NeurIPS, and ICLR),evaluating LLMs on two dimensions: High-Level planning completeness (main, ablation, and analysis experiments) and Low-Level configuration accuracy and rationality (datasets, baselines, and metrics). Benchmarking reveals three findings: (1) most LLMs cannot directly design high-quality experiments; (2) all LLMs exhibit a performance bottleneck in low-level configuration; and (3) search mode does not improve design quality. Furthermore, to address these challenges, we propose OptED, a novel agentic workflow to optimize LLM-based experimental design, that enhances LLM-based experimental planning through stage isolation, tool augmentation, and rule-based constraints, effectively alleviating the configuration bottleneck.
Comments32 pages, 7 figures