发表机构
University of Georgia; Emory University(佐治亚大学; 埃默里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Vestrum通过将执行失败转化为针对性的框架修改,在不训练任务模型的情况下提升智能体框架在多个基准上的性能,其核心是利用可迁移的失败类别和持久经验文件。
AI 中文摘要
智能体框架控制语言模型如何访问信息、使用工具、保留记忆并检查其工作。当每次评估都需要与环境进行长时间交互时,改进此类软件成本高昂。我们引入了Vestrum,一个框架,能够将执行轨迹中的失败转化为有针对性的框架修改,而无需训练任务模型。其核心假设是,同类任务可能表现出反复出现的失败,而这些失败的补救措施在该类任务内具有可迁移性。Vestrum将失败表示为可识别的类别,提出跨验证、检索、分解和知识综合的修改,并在将其作为整体评估之前筛选其适用范围。一个持久的经验文件为后续提议提供信息。在五个设置和两个基线框架中,冻结的框架提升了保留数据的性能:UltraHorizon从47.6提升至59.8(相对于GAM),Terminal-Bench 4 Hard在八个保留任务上通过检查的比例从63.7%提升至70.3%(相对于Claude Code),测试成本为1.03倍,细胞类型注释一致性在一个切片的保留部分上从67.5%提升至77.8%,同时在LoCoMo和AMA-Bench上也有提升。在我们的搜索中,基于证据的验证有助于中间步骤和最终答案,在中间步骤上成本更低,而被要求重建已完成答案的批评者破坏多于修复。在三个记忆基准上,Vestrum在每次配对评估中的得分也高于所评估的GEPA配置。
英文摘要
An agent harness controls how a language model accesses information, uses tools, preserves memory, and checks its work. Improving this software is costly when each evaluation requires a long interaction with an environment. We introduce Vestrum, a framework that turns failures in execution traces into scoped harness changes without training the task model. Its organizing overhypothesis is that tasks of a shared kind may exhibit recurring failures whose remedies transfer within that kind. Vestrum expresses failures as recognizable classes, proposes changes across verification, retrieval, decomposition, and knowledge synthesis, and screens their scope before evaluating them as a bundle. A persistent lessons file informs subsequent proposals. Across five settings and two baseline harnesses, the frozen harnesses improve held-out performance: UltraHorizon rises from 47.6 to 59.8 over GAM, Terminal-Bench 4 Hard from 63.7% to 70.3% of checks passed over Claude Code on eight held-out tasks at 1.03x test cost, and cell-type annotation agreement from 67.5% to 77.8% on held-out sections of one slide, alongside gains on LoCoMo and AMA-Bench. Across our searches, verification grounded in evidence helped both intermediate steps and final answers, at lower cost at intermediate steps, while critics asked to rebuild finished answers broke more than they repaired. On the three memory benchmarks, Vestrum also scores above the evaluated GEPA configurations in every paired evaluation.
Comments30 pages, 2 figures, 20 tables. Under review