使前瞻性记忆具备SLM形态:面向小模型智能体的类型化意图存储
Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents
浏览论文内容
中文总结 AI 辅助
本文针对小型智能体的前瞻性记忆任务,提出前瞻性意图存储(PIS)方法,在PM-Bench基准上实现集F1值的显著提升,使小型模型性能超越已发表的大型模型框架。
中文摘要 AI 辅助
前瞻性记忆是指在执行其他任务的同时,于未来恰当的提示下完成延迟意图的能力。目前基准测试已将其分离为智能体技能,但前沿大型语言模型(LLM)仍难以掌握:已发表的最佳PM-Bench框架仅达到65.1%的集F1值。本文认为该任务的核心是受模式约束的状态跟踪,而非开放式推理,当动作空间具备类型时,小型模型也可执行该任务。我们提出前瞻性意图存储(PIS),将生命周期逻辑置于代码中,将范围限定的语言工作交由模型处理。该框架具备智能体特性且无需训练:无需选择器微调,也无需轨迹蒸馏。在PM-Bench上,搭载PIS的DeepSeek-Chat达到82.9%的集F1值;对于Gemma-E2B,无存储时集F1仅为4.2%,在7种回溯记忆方法下最多为6.6%,而PIS可达到66.2%;在另一测试中,PIS达到70.1%的集F1值,而回溯记忆方法最多为54.4%。PIS在该基准上设定了新的当前最佳水平,且使小型模型能够超越已发表的大型模型框架。
英文摘要
Prospective memory means carrying out a deferred intention at the right future cue while other work continues. Benchmarks now isolate it as an agent skill, yet frontier LLMs still struggle: the best published PM-Bench scaffold reaches only 65.1% Set-F1. We argue that this loop is schema-constrained state tracking rather than open-ended reasoning, and that small models can execute it when the action space is typed. We propose the Prospective Intention Store (PIS) that puts lifecycle logic in code and scoped language work on the model. The scaffold is agentic and training-free: no selector fine-tuning and no trajectory distillation. On PM-Bench, DeepSeek-Chat with PIS reaches 82.9% Set-F1. On Gemma-E2B, Set-F1 is only 4.2% without a store and at most 6.6% under seven retrospective memories, while PIS reaches 66.2%. PIS further reaches 70.1% Set-F1, where retrospective memory methods stay at most 54.4%. PIS sets a new state of the art on this benchmark and enables small models to surpass the published large-model scaffold.
发表机构
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。