发表机构
Apple Inc.(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
智能语言模型系统面临上下文问题,本文提出共享选择性持久内存架构,识别保留四类可重用上下文,丢弃特定会话推理痕迹,实现共享协作重用。在企业场景和公共数据集实验中验证其有效性,提高任务完成率,降低成本,优于全历史持久化等极端方式。
AI 中文摘要
通过多轮工具使用生成代码的智能语言模型系统面临一个基本的上下文问题:每个会话从零开始,丢弃使先前会话高效的配置选择、域约束、数据模式和工具使用模式。天真地持久化整个对话历史在令牌使用上效率低下且适得其反:无关上下文会降低生成质量。我们引入了共享选择性持久内存,它能识别并保留四类可重用上下文(任务规范、数据模式、工具配置和输出约束),同时丢弃特定会话的推理痕迹。关键是,该内存是共享的:封装选择性内存的工作区可通过基于角色的访问控制在用户间转移,实现协作重用而无需冗余规范。我们在一个部署的协作工作区平台中实现了它,在该平台上,语言模型智能体从异构源(CSV、SQL、REST API和MCP服务器)生成、编辑和维护带git版本控制的工件(仪表板、报告和数据驱动文档)。一种互补的零令牌数据刷新机制将生成的程序与运行时数据解耦,实现工件重用而无需重新调用。在三个企业场景中,共享选择性持久内存实现了96%的任务完成率(无内存时为79%,全历史记录时为71%)。零令牌刷新消除了重复更新时对语言模型的重新调用(任务时间减少14倍),而摘要驱动生成使每次调用的令牌成本相较于原始数据注入降低了97倍。在四个公共数据集上的复制验证了其通用性,零令牌刷新在12次试验中全部成功。值得注意的是,天真的全历史持久化会因陈旧痕迹使智能体产生偏差,从而积极降低完成率,而选择性内存优于这两种极端情况。
英文摘要
Agentic LLM systems that generate code through multi-turn tool use face a fundamental context problem: each session starts from zero, discarding the domain constraints, data schemas, tool configurations, and output preferences that made previous sessions productive. We introduce shared selective persistent memory, an architecture that retains four categories of reusable context - task specifications, data schemas, tool configurations, and output constraints - while discarding session-specific reasoning traces, and that packages them into workspaces transferable across users under role-based access control. The resulting cost curve is non-monotonic. In a controlled replication on four public datasets, where a formatting specification is established once and then withheld, no memory completes 0/12 trials at 3.8K input tokens, selective memory completes 12/12 at 3.9K, and full conversation history completes 8/12 at 7.7K. What is kept matters more than how much is kept: the winning configuration costs essentially what the failing one does, and twice as much context does not improve on it. Both differences from no memory survive Bonferroni-corrected exact McNemar tests (p = 0.0005, p = 0.008); the two memory conditions separate on price rather than completion. We implement this in a deployed platform where agents produce git-versioned artifacts from CSV, SQL, REST, and MCP sources. A complementary zero-token data refresh contract decouples generated programs from runtime data, firing on 12/12 trials at a median 0.08s with no model call, while summary-driven data representation costs 97-431x fewer tokens than raw injection. Across 24 recurring enterprise tasks selective memory completes 23/24 against 19/24 and 17/24, though at that sample no pairwise difference reaches significance.
Comments11 pages, 2 figures, 4 tables