arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自设计评估器与热记忆用于长时程智能体

Self-Designed Evaluators and Warm Memory for Long-Horizon Agents

Saeid Asgari, Emre Kiciman, Leonardo de Oliveira Nunes, Ranveer Chandra

arXiv 2609.33717首次发表:更新:

发表机构

Microsoft(微软公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SelfSuite通过智能体自设计评估器与热记忆,在无奖励的长任务中实现自我评估与改进,性能优于无标签基线,并可与专家标签方法匹敌。

AI 中文摘要

一个在长任务流中部署的工具使用语言模型智能体无法获得奖励,因此它无法判断是否成功,无法安全地重试,也无法标记其改进所需的经验。我们提出了SelfSuite,在该方法中,智能体自身的基础模型,仅利用世界的公开材料,设计一个小型评估套件,包含加权评判器和基于任务的简报,将其冻结,并用它来门控保留最佳结果的重试,以及标记一个带类型、结果跟踪的记忆。在tau2-bench和AppWorld上匹配的五次重复基准测试中,SelfSuite在没有任何标签的情况下得分高于普通智能体,在tau2-bench上与给定十个专家标签的方法相匹配,而在AppWorld上落后于Agentic Context Engineering (ACE),后者通过代码执行提供直接的成功信号。在同一任务上进行的消融实验中,它在每次重复中都高于无标签的ACE,并且门控的第二次尝试是唯一在每次重复中移除都会造成性能下降的组件。我们还模拟了一个领域专家,对每个世界中的十个入职任务进行评分。使用这些标签来校准SelfSuite的评估器带来了小幅但一致的提升,而使用它们来预热ACE的记忆则使ACE与校准后的SelfSuite持平。在第二个模型系列上的单次运行研究显示了相同的排序。

英文摘要

A tool-using language-model agent deployed over a long stream of tasks receives no reward, so it cannot tell whether it succeeded, cannot safely retry, and cannot label the experience it needs to improve. We present SelfSuite, in which the agent's own base model, given only the world's public materials, designs a small evaluation suite of weighted judges and grounded per-task briefs, freezes it, and uses it to gate a keep-best retry and to label a typed, outcome-tracked memory. On matched five-repeat benchmarks over tau2-bench and AppWorld, SelfSuite scores above the plain agent without any labels, matches methods given ten expert labels on tau2-bench, and trails Agentic Context Engineering (ACE) on AppWorld, where code execution gives a direct success signal. In an ablation campaign run on the same tasks, it is above label-free ACE in every repeat, and the gated second attempt is the only component whose removal hurts in every repeat. We also simulate a subject-matter expert who grades ten onboarding tasks per world. Using those labels to calibrate SelfSuite's evaluator gives a small, consistent gain, and using them to warm up ACE's memory lifts ACE to tie calibrated SelfSuite. A single-run study on a second model family shows the same ordering.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑