arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体的肌肉记忆:编译而非仅检索

Muscle Memory for Agents: Compile not Merely Retrieve

Pouya Ghiasnezhad Omran, Soujanya Lanka, Qin Zhang, Tanya Dixit

arXiv 2608.08995首次发表:更新:

发表机构

Google Cloud FDE(谷歌云 FDE)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出将重复用户意图编译为专用专家智能体的肌肉记忆范式,通过四阶段流水线实现,在90个场景中专家触发时胜率达88.9%,优于传统检索模式,为智能体记忆设计提供新方向。

AI 中文摘要

大语言模型(LLM)智能体的记忆架构已趋于统一:将经验存储为文本、嵌入、反思或规则,在推理时进行检索,由通用编排器解读执行任务。本文认为该模式并非个性化的默认选择。我们提出肌肉记忆(Muscle Memory)这一概念,即把重复出现的用户意图编译为专用的专家智能体,将其定位为与检索不同的记忆范式,且编译更适用于当前助手给用户带来多轮负担的场景——用户需反复修正格式、深度和范围才能获得符合领域要求的答案。我们通过参考实现和经验证据支持这一观点。该实现是一个四阶段流水线(采集→分析→增强→评估),用于挖掘对话历史,分离行为模式与任务模式,输出经质量把关的可执行编译专家智能体,采用两阶段触发匹配。在5种用户角色的90个保留场景中,当专家智能体被触发时,增强型助手在36个案例中赢得32个,胜率达88.9%,个性化增益为+2.05,在1-4分尺度上的准确率损失仅为-0.28。我们还探讨了在此场景下编译为何比检索更合适、该结果对更广泛的记忆设计空间的启示以及剩余的开放问题。

英文摘要

Memory for LLM agents has converged on a single architectural pattern: store experience as text, embeddings, reflections, or rules; retrieve at inference time; let a general-purpose orchestrator interpret what to do. This paper argues that the pattern is the wrong default for personalization. We position Muscle Memory - the practice of compiling recurring user intent into purpose-built specialist agents - as a distinct memory paradigm from retrieval, and we argue that compilation is a better fit for the workloads where current assistants impose a multi-turn tax on their users: making them repeatedly correct format, depth, and scope to obtain a domain-appropriate answer. We support the position with a reference implementation and empirical evidence. The implementation is a four-phase pipeline (Harvest $\rightarrow$ Analyze $\rightarrow$ Augment $\rightarrow$ Evaluate) that mines conversational history, separates behavioral from task patterns, and emits quality-gated executable compiled specialists with two-stage trigger matching. On 90 held-out scenarios across five user personas, the augmented assistant wins 32 of 36 cases where a specialist fires, an 88.9% win rate, with a +2.05 personalization gain and only a $-0.28$ accuracy cost on a 1-4 scale. We discuss why compilation is better suited than retrieval in this regime, what the result implies for the broader memory design space, and what open problems remain.

Comments12 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑