MILO:通过编排式多智能体进化实现自动化工具链发现
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
浏览论文内容
中文总结 AI 辅助
提出MILO框架,通过多智能体协同进化自动发现智能体工具链,在多个基准上超越现有方法,显著提升性能并降低令牌消耗。
中文摘要 AI 辅助
现代智能体系统将人工智能模型与一个控制执行和环境交互的工具链相结合。工具链的设计对长周期性能影响显著,但其组合搜索空间巨大,需要大量人工投入,且随着模型更换必须重复进行。现有的自动化方法对该空间的探索较为狭窄,仅优化提示词或技能等组件,或陷入固定的、剥削性的搜索策略。我们提出了MILO(元进化岛屿编排),一个协同进化智能体工具链及其发现策略的框架。MILO结合了:(i)基于岛屿树的分层谱系记忆,利用被拒绝的变异作为负面证据;(ii)岛屿内变异智能体,利用全局搜索历史和父代特定反馈重写完整工具链;(iii)一个编排器,通过谱系嫁接和物种形成、变异智能体重分配和课程修订来适应搜索。在Terminal-Bench 2.1、PaperBench和DeepSWE上,使用前沿(Opus 4.8)和开放权重(gpt-oss-120b)模型,MILO发现的工具链优于八个最先进的工具链和六种搜索方法。使用Opus 4.8,MILO相对于其初始工具链的解决率分别提高了+12.0%、+28.3%和+10.3%,而先前最佳搜索增益分别为+4.5%、+18.3%和0%。在Terminal-Bench 2.1上,它达到86.1±2.0%,超过官方排行榜第一名(83.8±2.3%),同时比其初始工具链少使用26%的令牌。在EinsteinArena开放问题上,MILO改进了Erdős最小重叠问题的最佳已知上界(0.3808586→0.3808568),以及第一和第三自相关不等式(1.50274365→1.50274360;1.45081→1.44889)。
英文摘要
Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated methods explore this space narrowly, optimizing only components such as prompts or skills or becoming trapped by fixed, exploitative search strategies. We introduce MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them. MILO combines: (i) hierarchical lineage memory over island-based trees, using rejected mutations as negative evidence; (ii) per-island mutator agents that rewrite complete harnesses using global search history and parent-specific feedback; and (iii) an orchestrator that adapts search through lineage grafting and speciation, mutator reassignment and curriculum revision. Across Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models. With Opus 4.8, MILO improves resolution over its initial harness by $+12.0\%$, $+28.3\%$, and $+10.3\%$, respectively, compared with best prior-search gains of $+4.5\%$, $+18.3\%$, and $0\%$. On Terminal-Bench 2.1, it achieves $86.1 \pm 2.0\%$, exceeding the official leaderboard's top entry ($83.8 \pm 2.3\%$) while using 26\% fewer tokens than its initial harness. On EinsteinArena open problems, MILO improves best-known upper bounds for Erdős minimum-overlap ($0.3808586 \to 0.3808568$) and the first and third autocorrelation inequalities ($1.50274365 \to 1.50274360$; $1.45081 \to 1.44889$).
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
- AWS AI Labs(AWS AI 实验室)
- Carnegie Mellon University(卡内基梅隆大学)
- Washington University in St. Louis(圣路易斯华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。