arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40303cs.AI

一个强大的智能体进行自主机器学习工程需要多少约束框架?

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel Abbé

首次发表
浏览论文内容

中文总结 AI 辅助

本文发现,在相同时间预算和LLM骨干下,复杂约束框架对自主机器学习工程智能体无优势,骨干模型是性能主要驱动力,构建复杂机制回报甚微。

中文摘要 AI 辅助

近年来,自主机器学习工程(MLE)智能体在公开排行榜上取得了显著进展。由于在长周期任务中进展停滞以及大型语言模型(LLM)原语有限,现代MLE智能体被部署在日益复杂的机制之上:多智能体编排器、专用检索子智能体等。尽管此类约束框架不断扩展,但使用更原始但改进的编码智能体——其中LLM通过读、写和bash原语直接访问执行环境——在该领域却鲜受关注。在本文中,我们发现,在相同的时间预算和相同的前沿LLM骨干下,开源的最先进的约束框架相比单会话的最小约束编码智能体基线并无优势,这表明骨干是性能的主要驱动力。通过一系列大规模系统性消融研究,我们认为在编码智能体设置中,机制层变得冗余。我们得出结论,围绕强大模型精心构建约束框架的努力在当前MLE基准上回报甚微。

英文摘要

Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orchestrators, dedicated retrieval subagents, and more. While such harnesses expand, the use of more primitive but improved coding agents - where LLMs have direct access to the execution environment through read, write, and bash primitives - has received little attention in the field. In this paper we find that, under an equal time budget and the same frontier LLM backbone, open-source state-of-the-art harnesses provide no advantages over a single session of a minimal-harness coding agent baseline, pointing to the backbone as the primary driver for performance. Via a series of large-scale systematic ablation studies, we argue that the machinery layers become redundant in the coding agent setting. We conclude that the effort spent elaborating hand-crafted harnesses around strong models yields poor returns for current MLE benchmarks.

发表机构

  • EPFL(洛桑联邦理工学院)
  • Apple(苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

↑