AI4AI-Bench:用于递归自我改进算法设计中大语言模型智能体的基准测试
AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
浏览论文内容
中文总结 AI 辅助
本文提出AI4AI-Bench基准测试,测试大语言模型智能体设计训练算法的能力,实验显示现有智能体仅缩小了已有算法与最优值间不到五分之一的差距,发布相关资源供重复测量。
中文摘要 AI 辅助
递归自我改进(RSI)探究AI系统能否改进生成AI系统的过程,使下一个系统继承该改进,这一过程即训练算法:更优的目标或更新规则可提升后续每一轮运行的计算-能力交换率,包括生成下一个智能体的那一轮。因此RSI是否可行取决于智能体能否设计训练算法。目前没有基准测试能单独隔离这一能力:现有基准套件要么是通过收集数据,要么是通过调整超参数获胜,且都无法区分运行方式的改变与模型学习方式的改变。本文提出AI4AI-Bench,涵盖10个训练算法族的10个固定研究仓库。在每个任务中,智能体在1个B300上有4小时重写训练算法;其代码随后从头重新运行最多12小时,由对智能体隐藏的固定评估器,在相同流程下与仓库原始算法对比评分。由于10个指标不可比,每个任务被映射到一个刻度,其中0代表无信息模型,0.1代表仓库提供的算法,1.0代表任务最优值。在6个系统的29种配置下,所有10个任务的平均得分为0.166,最佳系统达到0.250:即使是最强的系统,也仅缩小了已有算法与最优值之间不到五分之一的差距。提交结果显示了该差距的去向:大多数提交完全未改变模型的学习方式,而确实改变的少数提交平均得分为0.226,其余为0.126。更多推理努力主要提升了尝试改变的意愿,使这类少数提交占比从8%升至64%,平均得分从0.094升至0.196。本文发布了该任务套件、评估器及所有已评分的提交,以便随着这些系统的变化可重复进行测量。
英文摘要
Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement. That process is the training algorithm: a better objective or update rule improves the compute\mbox{-}capability exchange rate for every subsequent run, including the one that produces the next agent. Whether RSI is feasible therefore turns on whether an agent can design training algorithms. No benchmark isolates that ability: existing suites are won by collecting data or by tuning hyperparameters, and none tells a change to how a run is executed apart from a change to how the model learns. We present AI4AI\mbox{-}Bench, 10 frozen research repositories spanning 10 training algorithm families. In each task, an agent has 4 hours on one B300 to rewrite the training algorithm; its code is then rerun from scratch for up to 12 hours and scored by a fixed evaluator hidden from the agent, against the repository's original algorithm under the same procedure. Because the 10 metrics are incommensurable, every task is mapped onto one scale on which $0$ is an uninformative model, $0.1$ is the algorithm the repository ships, and $1.0$ is the task optimum. Across 29 configurations of 6 systems on all 10 tasks the mean score is $0.166$, and the best system reaches $0.250$: even the strongest closes under a fifth of the distance between the algorithm that was already there and the optimum. The submissions show where that distance went: most never change how the model learns at all, and the minority that do average $0.226$ against $0.126$ for the rest. More reasoning effort mostly buys the willingness to go there, taking that minority from $8\%$ of submissions to $64\%$ and the mean score from $0.094$ to $0.196$. We release the task suite, the evaluators and every scored submission, so that the measurement can be repeated as these systems change.
发表机构
- Navers Lab(Navers实验室)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。