arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

即时记忆:为LLM智能体学习策展任务自适应记忆

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

Yefan Zhou, Yang Li, Zeyu Leo Liu, Semih Yavuz, Shafiq Joty

arXiv 2609.27334首次发表:更新:

AI 中文总结

针对现有记忆系统在写入时固化摘要、无法适应未来查询的问题,提出即时记忆(JitMem),在读取时基于当前任务策展原始轨迹,生成任务自适应载荷,并在ALFWorld、WebShop和τ²-bench上超越基线,显著提升成功率。

AI 中文摘要

智能体记忆系统通过复用过往经验来提升未来表现,然而现有大多数设计在写入时进行记忆策展:一旦任务完成,其轨迹便被提炼为固定产物,如反思、工作流、技能或推理策略,之后通过相似性检索。这迫使系统在未知未来查询之前就决定哪些内容值得记住,不可逆地丢弃信息,并生成与查询无关的摘要,该摘要必须服务于许多可能的下游任务。学习这样的写入时策展者也十分困难,因为存储决策的价值可能只有在相关查询到来时才显现,而该查询可能发生在许多任务之后,从而形成长时程信用分配问题。我们转而保留原始轨迹,并将策展推迟到读取时,即当前任务已知之时。给定检索到的轨迹和新任务,记忆策展者会合成一个紧凑的、任务自适应的载荷,以满足即时需求。由于该载荷在同一任务上被消费,策展者可以直接从即时任务成功中训练,避免了延迟的效用信号以及人为分组相关任务的需要。在ALFWorld、WebShop和τ²-bench上,我们的即时记忆(JitMem)持续优于无记忆智能体以及启发式和学习的写入时记忆方法,分别比最强基线提高了16.2、16.3和3.9个绝对成功率百分点。值得注意的是,即使未经训练的策展者已能与这些基线竞争甚至超越它们,这表明任务自适应的读取时策展本身就是提升的主要来源;而训练策展者则进一步放大了改进效果。

英文摘要

Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This forces the system to decide what is worth remembering before the future query is known, irreversibly discarding information and producing a query-independent summary that must serve many possible downstream tasks. Learning such a write-time curator is also difficult because the value of a storage decision may only become apparent when a relevant query arrives, potentially many tasks later, creating a long-horizon credit-assignment problem. We instead retain raw trajectories and defer curation until read time, when the current task is known. Given the retrieved traces and the new task, a memory curator synthesizes a compact, task-adaptive payload tailored to the immediate need. Because this payload is consumed on the same task, the curator can be trained directly from immediate task success, avoiding delayed utility signals and the need to artificially group related tasks. Across ALFWorld, WebShop, and $τ^2$-bench, our Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively. Notably, even an untrained curator is already competitive with or surpasses these baselines, showing that task-adaptive read-time curation itself is a major source of the gain; training the curator further compounds the improvement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑