发表机构
University of Virginia; Microsoft(弗吉尼亚大学; 微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体进化中求解器与评估器相互耦合的挑战,提出DUET框架,通过联合优化求解器与评分器智能体,在四个基准上同时提升两者性能并超越固定评分器的基线。
AI 中文摘要
智能体工作流正越来越多地应用于技术、金融和企业运营等领域。随着这些智能体被广泛部署,持续改进它们变得日益重要。这带来了一个直接的挑战:智能体应如何进化?这种进化需要有效的评估,能够评估结果并为优化提供有用的反馈。随着智能体的进化,其行为和失败模式也可能发生变化,使得固定的评估器越来越不适用。另一个基本问题是:我们应如何评估一个正在进化的智能体?这两个挑战本质上是耦合的;智能体行为的变化可能暴露当前评估器的局限性,而更强的评估器则能为改进智能体提供更具信息量的反馈。受这种交互作用的启发,我们提出了DUET,一个联合优化求解器智能体和评分器智能体以同时改进两者的框架。DUET迭代地选择训练任务,用求解器执行这些任务,用评分器评估产生的结果,并使用一个工具使用的更新模块来修订求解器和评分器,在数轮中交替进行。通过在优化循环内更新评分器,DUET将评估从固定的反馈来源转变为与求解器一同适应的首要优化目标。在四个智能体基准上的实验表明,DUET同时提高了求解器和评分器的性能,并且始终优于使用固定评分器优化求解器的基线方法。
英文摘要
Agentic workflows are increasingly used across domains such as technology, finance, and enterprise operations. As these agents become more widely deployed, continually improving them becomes increasingly important. This raises an immediate challenge: How should the agent evolve? This evolution requires effective evaluation that can assess outcomes and provide useful feedback for optimization. As the agent evolves, its behaviors and failure modes may also change, making a fixed evaluator increasingly inadequate. Another fundamental question: How should we evaluate an evolving agent? These two challenges are inherently coupled; changes in agent behavior can expose limitations of the current evaluator, while a stronger evaluator provides more informative feedback for improving the agent. Motivated by this interaction, we introduce DUET, a framework that jointly optimizes a solver agent and a grader agent to improve both. DUET iteratively selects training tasks, executes them with the solver, evaluates the resulting outcomes with the grader, and uses a tool-using update module to revise the solver and the grader, alternating between the two across rounds. By updating the grader within the optimization loop, DUET turns evaluation from a fixed source of feedback into a first-class optimization objective that adapts alongside the solver. Experiments across four agent benchmarks show that DUET improves both solver and grader performance and consistently outperforms baselines that optimize the solver with a fixed grader.