arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MESH-Harness:通过波段引导的组合进化实现自我改进的智能体框架

MESH-Harness: Self-Improving Agent Harnesses via Bandit-Guided Compositional Evolution

Zhiwei Shang, Yu Huo, Mingrong Gong, An Yan, Zikun Qu, Junhao Dong, Bryan Kian Hsiang Low, Chenglin Wu, Zhongxiang Dai

arXiv 2610.05300首次发表:更新:

发表机构

The Chinese University of Hong Kong, Shenzhen; The Chinese University of Hong Kong; Nanyang Technological University; Fudan University; National University of Singapore; DeepWisdom(香港中文大学(深圳); 香港中文大学; 南洋理工大学; 复旦大学; 新加坡国立大学; 深度智慧)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MESH-Harness通过模块化组合与LinUCB引导的进化搜索,在固定模型权重和有限预算下系统改进智能体框架,在多项任务上超越基线并降低总成本。

AI 中文摘要

智能体框架(agent harness)是组织上下文、维护状态并协调语言模型工具调用的代码。我们研究了在有限的评估预算下,保持模型权重不变的同时如何改进框架。我们的方法MESH-Harness将每个框架组织为具有明确角色特定接口的功能模块,允许每个模块的替代实现被替换和重新组合。它使用共享模块表示和全协方差LinUCB,基于预测性能和探索价值对候选组合进行评分。混合起点坐标上升法选择完整的配置进行评估,而无需枚举组合空间。验证轨迹随后指导局部代码编辑,生成的候选被纳入固定容量的角色特定池中以供后续重组。在文本任务、检索增强的数学推理、代码生成和交互式科学任务上,在匹配的候选评估预算下,MESH-Harness分别比Meta-Harness高出5.70、7.01、2.00和5.00分。迭代框架优化使MESH-Harness比其首轮配置提高5.63-7.79分。对于所报告的配置,总测试时成本比Meta-Harness低44.2%,而包括搜索在内的总成本低14.6%。这些结果表明,将模块级设计复用与反馈驱动的组合搜索相结合,可以在控制总体优化成本的同时系统地改进智能体框架。

英文摘要

An agent harness is the code that organizes context, maintains state, and coordinates tool calls for a language model. We study how to improve the harness under a limited evaluation budget while keeping model weights fixed. Our method, MESH-Harness, organizes each harness into functional modules with explicit role-specific interfaces, allowing alternative implementations of each module to be substituted and recombined. It uses shared module representations and full-covariance LinUCB to score candidate combinations based on predicted performance and exploration value. Mixed-start coordinate ascent selects complete configurations for evaluation without enumerating the combinatorial space. Validation traces then guide local code edits, and the resulting candidates are incorporated into fixed-capacity role-specific pools for subsequent recombination. On text tasks, retrieval-augmented mathematical reasoning, code generation, and interactive scientific tasks, MESH-Harness outperforms Meta-Harness by 5.70, 7.01, 2.00, and 5.00 points, respectively, under matched candidate-evaluation budgets. Iterative harness optimization improves MESH-Harness by 5.63-7.79 points over its first-round configurations. For the reported configurations, aggregate test-time cost is 44.2% lower than that of Meta-Harness, while total cost including search is 14.6% lower. These results show that combining module-level design reuse with feedback-driven compositional search can systematically improve agent harnesses while keeping overall optimization cost under control.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑