发表机构
M365; Microsoft Research(M365; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出伪自蒸馏框架PSD,让小型语言模型通过提示词从黑盒预言机蒸馏记忆构建能力,以低成本达到甚至超越大模型性能,并具备分布外泛化能力。
AI 中文摘要
记忆系统正成为LLM智能体的核心组件,但构建和维护记忆仍然代价高昂,因为它依赖于对大型专有语言模型的重复调用。这一成本为大规模部署记忆增强型智能体设置了主要障碍。在本文中,我们提出了伪自蒸馏(PSD),这是一个通过多阶段训练流程,从强大的黑盒预言机中蒸馏行为,使小型语言模型(SLM)能够构建层次化记忆表征的框架。标准蒸馏方法需要访问教师模型的logits或隐藏状态,而封闭模型不暴露这些信息。与传统的自蒸馏设置(其中监督信号来自模型自身的预测、采样轨迹或聚合输出)不同,PSD在通过提示词引入外部预言机知识的同时,实现了单模型蒸馏设置。PSD使用单个小模型扮演两个角色:一个教师,它看到包含预言机答案作为参考上下文的特权提示;以及一个学生,它只看到任务提示。学生学会复现教师的输出分布,将预言机引导的行为吸收到自身权重中,而无需访问预言机的内部信息。在LoCoMo上,经过PSD训练的Qwen3-0.6B、1.7B和4B模型在下游检索任务上以极低的部署成本匹配或超过了GPT-4.1-mini,其中离策略PSD在大多数条件下取得了最强结果。我们进一步表明,这种记忆构建能力可以分布外迁移到LongMemEval,尽管学生仅在LoCoMo上训练,未接触过LongMemEval数据。
英文摘要
Memory systems are becoming a core component of LLM agents, but constructing and maintaining memory remains expensive because it relies on repeated calls to large proprietary language models. This cost creates a major barrier to deploying memory-enhanced agents at scale. In this paper, we present Pseudo Self-Distillation (PSD), a framework that enables small language models (SLMs) to construct hierarchical memory representations by distilling behavior from a strong black-box oracle through a multi-stage training pipeline. Standard distillation methods require access to teacher logits or hidden states, which closed models do not expose. Unlike conventional self-distillation settings, where supervision is derived from a model's own predictions, sampled rollouts, or aggregated outputs, PSD enables a single-model distillation setup while channeling external oracle knowledge through the prompt. PSD uses a single small model in two roles: a teacher that sees a privileged prompt containing the oracle's answer as reference context, and a student that sees only the task prompt. The student learns to reproduce the teacher's output distribution, absorbing oracle-guided behavior into its own weights without accessing the oracle's internals. On LoCoMo, PSD-trained Qwen3-0.6B, 1.7B, and 4B match or exceed GPT-4.1-mini on downstream retrieval at a fraction of the deployment cost, with off-policy PSD achieving the strongest results across most conditions. We further show that this memory-construction capability transfers out-of-distribution to LongMemEval, despite the students being trained exclusively on LoCoMo with no exposure to LongMemEval data.