发表机构
University of Cambridge; Tencent(剑桥大学; 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对现有智能体设计方法的缺陷,提出以问题为中心的优化方案,构建ADIAS框架,在五个交互式基准中较最强基线平均提升25.2%,消融实验验证了关键机制的有效性。
AI 中文摘要
自动化智能体设计通过迭代修订、评估和反馈摘要来改进智能体框架。现有方法大多以候选为中心:跨轮次经验围绕候选智能体组织,这使得修复过程隐式存在,导致修复目标效率低下、部分进展整合缓慢,以及无效干预措施在多轮次中传播。因此,我们提出以问题为中心的智能体优化,将修复进展作为显式持久问题状态推进,以指导优化,而非每轮从候选历史中重新推导。我们在ADIAS(一种用于自动化全代码智能体设计的框架)中实例化该方案,包含两种机制:持久问题状态维持稳定的问题标识、生命周期状态、支持证据及干预结果历史;以问题为中心的优化利用该状态联合提出后续聚焦全代码修改的修复目标和修订方向。在五个交互式基准测试中,ADIAS平均比最强基线高出25.2%,且在四个主干模型上均实现一致增益。受控消融实验进一步表明,移除持久问题状态或用候选中心策略替换以问题为中心的修订,会导致性能下降高达40.7%。
英文摘要
Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents, which leaves the repair progress implicit. This causes inefficient repair targeting, slow consolidation of partial progress, and propagation of ineffective interventions across rounds. Therefore, we formulate issue-centric agent optimization, in which repair progress is carried forward as an explicit persistent issue state to guide optimization, rather than re-derived from candidate history in each round. We instantiate the formulation in ADIAS, a framework for automated full-code agent design with two mechanisms. A persistent issue state maintains stable issue identities, lifecycle status, supporting evidence, and intervention-outcome histories. Issue-guided optimization uses this state to jointly propose repair targets and revision directions for subsequent focused full-code modification. Across five interactive benchmarks, ADIAS outperforms the strongest baseline by 25.2% on average and achieves consistent gains across four backbone models. Controlled ablations further show that removing persistent issue state or replacing issue-centric revision with candidate-centric policies leads to performance drops of up to 40.7%.
Comments23 pages, 7 tables, 5 figures