arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

循环中的记忆:进程内检索作为语言智能体的扩展工作记忆

Memory in the Loop: In-Process Retrieval as Extended Working Memory for Language Agents

Yusuf Khan, Carlo Lipizzi

arXiv 2607.05690首次发表:更新:

发表机构

Stevens(史蒂文斯理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究语言智能体中记忆进入循环的情况,通过对比网络存储和进程内存储的延迟,利用扩展思维理论提出进程内检索可作扩展工作记忆,经实验验证能提升召回率,还重新定位了操作瓶颈。

AI 中文摘要

语言智能体运行一个循环——观察、推理、行动——但它们推理所依据的记忆位于循环之外:一个存储库每轮最多被查询一次。我们研究记忆进入循环的情况,即在每一步进行读写。一直以来的障碍是延迟:联网存储库的响应时间为数十到数百毫秒,当检索成本高昂时,进程内检索会使端到端延迟增加多达83倍。先前的工作是管理这种成本而非质疑它:服务层调度隐藏了延迟,“内存优先”设计将检索限制为每轮一次。我们认为延迟是存储库所在位置的属性,而非循环内模式的属性:进程内存储库的响应时间约为100微秒,比网络模式低三个数量级,以这种速度,每步的成本可以忽略不计。根据扩展思维理论的等效原则,一个足够快且能持续直接可用的存储库就成为了扩展工作记忆,而不仅仅是智能体咨询的工具。前提是因果关系:保持每轮固定的内存延迟预算,仅改变存储库的响应速度,冗余动作会随着延迟单调增加——进程内速度下12次中有0次,110毫秒云往返时(gpt - 5 - 纳米,gpt - 5 - 迷你;精确排列p = 0.0079)12次中有7.2次。我们端到端地展示了这种情况:在有界窗口下的四个GPT - 5类模型中,使用循环内记忆时召回率从0/5提高到3.6 - 4.8/5,存储操作的p50为80 - 165微秒——尽管一个按指令每次回复都复述的基线也能完美解决问题,但令牌成本会随着工作集增加。在任何运行中存储库都从未丢失过事实(244次写入中有244次保留);每次未命中都可追溯到智能体的读取策略,而非存储库。我们的测量还重新定位了瓶颈:每步的主要成本是嵌入(通过网络约200 - 400毫秒);将进程内存储库与一个小型本地嵌入器配对可使整个操作回到约40微秒的测量值。

英文摘要

Language agents run a loop - observe, reason, act - but the memory they reason over sits outside it: a store queried at most once per turn. We study the regime where memory moves inside the loop, read and written on every step. The obstacle has always been latency: networked stores answer in tens to hundreds of milliseconds, and in-loop retrieval can inflate end-to-end latency by up to 83x when retrieval is expensive. Prior work manages that cost rather than questioning it: serving-layer scheduling hides it, "memory-first" designs ration retrieval to once per turn. We argue latency is a property of where the store lives, not the in-loop pattern: an in-process store answers in ~100us, three orders of magnitude below the network regime, and at that speed the per-step tax collapses. By the extended-mind thesis's parity principle, a store fast enough to be constantly and directly available becomes extended working memory, not a tool the agent merely consults. The premise is causal: holding a fixed per-turn memory-latency budget and varying only the store's answer speed, redundant actions rise monotonically with latency - 0.0 of 12 at in-process speed, 7.2 of 12 at a 110ms cloud round trip (gpt-5-nano, gpt-5-mini; exact permutation p=0.0079). We demonstrate the regime end-to-end: across four GPT-5-class models under a bounded window, recall improves from 0/5 to 3.6-4.8/5 with in-loop memory, store ops at p50 80-165us - though an instructed restate-every-reply baseline also solves it perfectly, at a token cost that grows with the working set. The store never lost a fact in any run (244 of 244 writes kept); every miss traces to the agent's read policy, not the store. Our measurements also relocate the bottleneck: the dominant per-step cost is embedding (~200-400ms over the network); pairing the in-process store with a small local embedder returns the complete operation to a measured ~40us.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑