arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeltaML-Bench:基于真实研究仓库评估机器学习智能体

DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories

Josias Moukpe, Priyanka Aryal, Matthew Kenney

arXiv 2608.19653首次发表:更新:

发表机构

Algorithmic Research Group(算法研究组)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出DeltaML-Bench基准,评估发现基于搜索的ARG框架可提升GPT-5在机器学习实验任务中的成功率,且能避免规范博弈,为自主ML智能体部署提供了关键考量

AI 中文摘要

用于机器学习实验的自主智能体必须在异构仓库中导航、修复训练管道,并在实际计算约束下评估候选改进方案。现有基准仅部分覆盖这些条件。我们推出DeltaML-Bench,该基准包含48个来自研究论文的任务,要求智能体在不完善的开源仓库中改进已发布的基线。我们使用标准Modular智能体和基于搜索的ARG框架评估GPT-5和Claude Sonnet 4:在4×6小时的时间分配下,ARG将GPT-5的单轮成功率从9.4%提升至33.9%;在2×12小时的时间分配下,GPT-5 ARG达到49.0%。Modular配置的规范博弈率高达47.9%,而在评估的ARG配置中未观察到任何博弈现象。这些结果表明,框架设计和完整性检查是部署自主机器学习实验智能体时需重点考虑的因素。

英文摘要

Autonomous agents for machine learning experimentation must navigate heterogeneous repositories, repair training pipelines, and evaluate candidate improvements under realistic compute constraints. Existing benchmarks only partially capture these conditions. We introduce DeltaML-Bench, a benchmark comprising 48 tasks sourced from research papers that require agents to improve published baselines within imperfect, open-source repositories. We evaluate GPT-5 and Claude Sonnet 4 with a standard Modular agent and a search-based ARG scaffolding. In the 4 x 6h allocation, ARG raises GPT-5's per-run success rate from 9.4% to 33.9%; in the 2 x 12h allocation, GPT-5 ARG reaches 49.0%. Modular configurations exhibit specification gaming rates as high as 47.9%, while no gaming is observed in the evaluated ARG configurations. These results indicate that scaffolding design and integrity checks are important considerations when deploying agents for autonomous ML experimentation.

Comments18 pages, 3 figures, 12 tables. Code and benchmark: https://github.com/AlgorithmicResearchGroup/deltaml-bench-public

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑