DAEDALUS:从自生成任务中引导智能体记忆
DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks
浏览论文内容
中文总结 AI 辅助
DAEDALUS通过探索者与求解器配对自生成任务,从失败中推导启发式规则并构建可复用记忆,无需人工指南或预言机验证器,在多个基准上显著提升成功率并降低推理成本。
中文摘要 AI 辅助
大语言模型智能体在新环境中往往缺乏可靠行动的操作性知识,因为它们必须自行发现特定的工具行为或环境约定。由于没有对过往尝试的记忆,它们会在不同任务中重复同样的错误,导致更多的任务失败和更长的轨迹。为解决这一问题,智能体系统通常依赖人工编写的指南或基于训练任务和预言机验证器构建的程序性记忆,而这两者都需要对环境的先验知识。我们提出DAEDALUS,一种无需现有任务或预言机验证器、从自生成练习中引导可复用智能体记忆的方法。DAEDALUS配对两个智能体:一个探索者,与环境交互以生成具有挑战性但可解决的任务;一个求解器,尝试解决这些任务。从每次求解器失败中推导出一条启发式规则,并且仅当求解器在上下文中反复使用该规则成功后才被接受。这些结果也为探索者提供反馈,以调整未来任务的难度。被接受的启发式规则随后被整合到一个记忆库中,供测试时使用。在AppWorld、τ²-bench和AutomationBench上,与无记忆基线相比,DAEDALUS将平均成功率提高了最多15.9个百分点,pass^5提高了最多2.2倍,并且与使用训练任务的方法竞争力相当,而推理成本低于大多数方法。我们表明,性能提升在较小的探索预算下即可显现,并且其启发式规则也惠及其他模型家族的智能体。我们的消融研究进一步揭示,求解器轨迹是推导有效启发式规则所需的关键信息,而将早期发现分解可使探索更具成本效益。除记忆构建外,我们发现DAEDALUS生成的任务在按性能对模型排序时可作为基准任务的代理。代码和工件:此HTTP URL。
英文摘要
LLM agents often lack the operational knowledge to act reliably in new environments, as they must discover specific tool behaviors or environment conventions on their own. Without memory of past attempts, they repeat the same mistakes across tasks, leading to more task failures and longer trajectories. To address this, agentic systems typically rely on human-written guidelines or on procedural memory built from training tasks and an oracle verifier, both of which require prior knowledge of the environment. We present DAEDALUS, a method for bootstrapping reusable agent memory from self-generated practice without existing tasks or oracle verifiers. DAEDALUS pairs two agents: an explorer that interacts with the environment to generate challenging yet solvable tasks, and a solver that attempts them. A heuristic is derived from each solver failure and accepted only after the solver repeatedly succeeds with that heuristic in context. These outcomes also provide feedback for the explorer to refine the difficulty of future tasks. Accepted heuristics are then consolidated into a memory bank for test-time use. Across AppWorld, $τ^2$-bench, and AutomationBench, DAEDALUS improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, and is competitive with methods using training tasks, at a lower inference cost than most. We show that performance gains already emerge with a small exploration budget, and that its heuristics also benefit agents from other model families. Our ablations further reveal that solver traces provide the key information needed to derive effective heuristics, while factorizing early discoveries makes exploration more cost-efficient. Beyond memory construction, we find that the tasks generated by DAEDALUS can serve as a proxy for benchmark tasks when ranking models by performance. Code and artifacts: www.github.com/illuin-tech/daedalus.
发表机构
- Illuin Technology
机构由 AI 辅助整理,请以论文原文为准。