arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07169cs.AIcs.LG

智能体记忆蒸馏:用分层教师记忆赋能小型大语言模型智能体

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

  • KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

Taeil Kim, Kangsan Kim, Sung Ju Hwang

AI总结:

本研究提出无需训练的Agent Memory Distillation框架,通过从大型教师智能体构建三类分层记忆迁移知识,提升小型大语言模型智能体的工具使用性能,在三个基准上实现显著准确率提升且优于现有基线。

AI中文摘要:

记忆系统已被证实有望提升智能体性能,但其在小型语言模型中的潜力仍未得到充分探索,小型语言模型难以自行生成足够多的成功轨迹。我们提出Agent Memory Distillation(AMD,智能体记忆蒸馏),这是一种无需训练的框架,通过分层记忆将大型教师智能体的结构化知识迁移至小型学生智能体。AMD从教师的成功轨迹中构建三种互补的记忆类型:Workflow memory(工作流记忆)编码任务级策略,Subtask memory(子任务记忆)提供中间粒度的具体行为示例,Function memory(函数记忆)捕获每个函数的调用约定和常见陷阱。工作流记忆与子任务记忆在每个任务开始时主动注入,而函数记忆则在工具调用错误时被动检索。我们在三个工具使用基准上使用四个学生模型(参数规模为4B至8B),以GPT-5-mini作为教师评估AMD,在AppWorld、BFCL V3和ToolSandbox上分别实现了27.2个百分点、11.2个百分点和3.4个百分点的平均准确率提升,且始终优于现有的基于记忆的基线方法。进一步分析表明,子任务记忆贡献的提升最大,教师的有效性取决于教师能力和学生兼容性,4B规模的学生从AMD中受益最多。

英文摘要:

Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which struggle to generate sufficient successful trajectories on their own. We propose Agent Memory Distillation (AMD), a training-free framework that transfers structured knowledge from a large teacher agent to a small student agent through hierarchical memory. AMD constructs three complementary memory types from successful teacher trajectories: Workflow memory encodes task-level strategies, Subtask memory provides concrete behavioral examples at an intermediate granularity, and Function memory captures per-function calling conventions and common pitfalls. Workflow and Subtask memories are injected proactively at the start of each task, while Function memory is retrieved reactively upon tool-calling errors. We evaluate AMD on three tool-use benchmarks using four student models (4B-8B parameters) with GPT-5-mini as the teacher, achieving average accuracy gains of 27.2%p, 11.2%p, and 3.4%p on AppWorld, BFCL V3, and ToolSandbox, while consistently outperforming existing memory-based baselines. Further analysis shows that Subtask memory contributes the largest gains, teacher effectiveness depends on both teacher capability and student compatibility, and 4B-sized students benefit most from AMD.

补充信息

↑