arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28003cs.LG

从失败中学习:面向小型语言模型工具使用智能体的异构图记忆

Learning from Failures: Heterogeneous Graph Memory for Small Language Model Tool-Using Agents

Jiaxing Li, Lei Song, Rui Dong, Youyong Kong

首次发表
浏览论文内容

中文总结 AI 辅助

针对中小语言模型工具智能体在长时程任务中的结构性错误,提出基于异构图记忆的失败感知检索框架FRESH,通过结构化经验提升任务成功率与工具使用可靠性。

中文摘要 AI 辅助

中小型语言模型为工具使用智能体提供了经济高效的执行器,使其在本地和大规模部署中具有吸引力。然而,在长时程且具有状态的环境下,它们常常犯下结构性错误,例如遗漏必要的观察结果、过早执行写入操作、重复失败的调用以及违反动作前置条件。这些错误可能导致状态更新不正确、策略违规以及代价高昂或不可逆的后果,使得可靠的工具执行成为一项关键的部署挑战。现有的微调方法需要大量的数据和计算资源,而扁平记忆可能检索到失败的动作,却未能保留其因果上下文或安全条件。在本文中,我们提出了FRESH,一种基于经验结构异构图(Experience-Structured Heterogeneous graphs)的失败感知检索框架(Failure-aware Retrieval framework),它将历史成功与失败转化为工具使用智能体的结构化外部经验。通过显式建模任务、动作、错误、修复和执行条件之间的依赖关系,FRESH帮助冻结的语言模型复用可靠策略、避免重复失败,并在有状态工具交互中做出更安全的决策。在τ-Bench和AppWorld上使用多个开源模型进行的实验表明,与无记忆智能体及具有代表性的基于记忆的基线方法相比,FRESH持续提升了任务成功率和工具使用可靠性。

英文摘要

Small and medium-sized language models offer cost-effective executors for tool-using agents, making them attractive for local and large-scale deployment. However, in long-horizon and stateful environments, they often make structural errors such as missing required observations, performing premature writes, repeating failed calls, and violating action preconditions. These errors can lead to incorrect state updates, policy violations, and costly or irreversible consequences, making reliable tool execution a critical deployment challenge. Existing fine-tuning approaches require substantial data and computation, while flat memory may retrieve failed actions without preserving their causal context or safety conditions. In this paper, we propose FRESH, a Failure-aware Retrieval framework over Experience-Structured Heterogeneous graphs, which transforms historical successes and failures into structured external experience for tool-using agents. By explicitly modeling the dependencies among tasks, actions, errors, repairs, and execution conditions, FRESH helps frozen language models reuse reliable strategies, avoid recurring failures, and make safer decisions in stateful tool interactions. Experiments on $τ$-Bench and AppWorld with multiple open-source models show that FRESH consistently improves task success and tool-use reliability over no-memory agents and representative memory-based baselines.

发表机构

  • Southeast University(东南大学)

机构由 AI 辅助整理,请以论文原文为准。

↑