arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自改进分支混合体用于智能体工具优化

Mixture of Self-Improving Branches For Agent Harness Optimization

Haoyu Dong, Yuhang Zhou, Zihao Lin, Yifan Wu, Bo Peng, Mingyi Wang, Xiangjun Fan, Lizhu Zhang, Zhuokai Zhao

arXiv 2609.37834首次发表:更新:

发表机构

Meta; Duke University; University of California, Davis(Meta; 杜克大学; 加州大学戴维斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对工具优化中固定开发集和提议策略导致局部最优的问题,提出分支化自适应搜索,通过进化开发子集和提议策略生成互补工具,并利用路由器选择分支头,在数学推理和编码基准上超越Meta-Harness达34.8%。

AI 中文摘要

工具优化为递归自我改进(RSI)提供了一个实用场景,其中智能体生成的修改通过执行反馈影响后续更改。最近的工作如Meta-Harness通过迭代代码生成和评估实现了这一过程,但保留了固定的开发集和提议策略。这些约束将进化引导到单一搜索轨迹上,增加了收敛到局部最优的风险。我们通过将搜索组织为具有进化开发子集和提议策略的分支,使改进过程本身具有适应性。每个分支保留由其领先工具比其他分支的工具解决更多开发案例,丢弃所有分支的领先工具都解决的案例,并使用其自身的搜索历史修订其提议策略。为了部署由此产生的互补工具,我们提出一个路由器,在执行前为每个新输入选择一个开发选定的分支头。在数学推理和智能体编码基准上,我们的系统相对于Meta-Harness在奥林匹克级数学推理上实现了34.8%的相对改进,在Terminal-Bench 2.0上实现了11.6%,在SWE-bench Lite上实现了3.8%,工具选择和路由器配置仅基于开发数据。这些结果表明,进化分支目标和提议策略可以产生互补的工具,路由器无需访问测试结果即可结合其优势。

英文摘要

Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this process through iterative code generation and evaluation, but retains a fixed development set and proposal policy. These constraints channel evolution along a single search trajectory, increasing the risk of converging to a local optimum. We make the improvement process itself adaptive by organizing search into branches with evolving development subsets and proposal policies. Each branch retains development cases solved by more of its leading harnesses than by those of other branches, drops cases solved by every leading harness across all branches, and revises its proposal policy using its own search history. To deploy the resulting complementary harnesses, we propose a router to select one development-selected branch head for each new input before execution. Across mathematical reasoning and agentic coding benchmarks, our system achieves relative improvements over Meta-Harness of 34.8% on Olympiad-level mathematical reasoning, 11.6% on Terminal-Bench 2.0, and 3.8% on SWE-bench Lite, with harness selection and router configuration based solely on development data. These results show that evolving branch objectives and proposal policies can yield complementary harnesses whose strengths a router combines without access to test outcomes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑