arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26637cs.CLcs.AI

面向LLM智能体的基于文件系统的内存:组织、演化与可持续性

Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability

Sizhe Zhou, Sheldon Yu, Hui Wei, Junda Wu, Siru Ouyang, Yizhu Jiao, Shijia Pan, Julian McAuley, Yu Zhang, Tong Yu, Jiawei Han

首次发表
浏览论文内容

中文总结 AI 辅助

本研究首次系统探索面向LLM智能体的基于文件系统的内存,定义三个智能体角色开展实验,发现有组织存储可减半检索成本,但多数智能体组织会退化,工具集对存储形态的影响与更换模型相当,将文件系统设为智能体内存的设计空间。

中文摘要 AI 辅助

部署的LLM智能体越来越多地将其长期内存存储为文件系统:一个由智能体自身通过通用文件工具读取、写入和重组的markdown文件目录树。然而,研究在很大程度上忽略了这种媒介:先前的系统设计定制了内存表示并研究了其上的检索,却未验证默认设置的两个工作假设:即智能体可在记忆积累、冲突和失效时将不断增长的存储组织起来,且这种组织是有益的。我们首次系统探索面向LLM智能体的基于文件系统的内存,将该设置形式化为围绕一个内存文件系统的三个角色:管理智能体整合并组织传入内容,搜索智能体用引用的源回答查询,执行智能体提供任务轨迹并提炼为技能,将陈述性记忆和技能统一到单个存储中。在长对话基准和具身任务中,我们改变内存形态(智能体组织的层级、逐字转储、分块检索)、流规模、工具 harness(沙箱shell、内存工具式函数、各类搜索工具)以及管理和搜索智能体的能力,跟踪内存增长时的答案质量、成本和存储健康状况。有组织的存储确实能带来搜索经济性:当内容量大时,有组织的存储的检索成本约减半。但如今的智能体未达到默认设置的承诺:在我们的增长研究中,除最强管理智能体外,所有智能体的组织都出现退化,且我们测量的智能体均未将组织本身转化为更好的答案。模型并非影响存储形态的唯一因素:仅改变工具集对存储形态的影响与更换模型一样大。本研究将文件系统默认设置从一个假设转变为智能体内存的设计空间。

英文摘要

Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that the agent itself reads, writes, and reorganizes through generic file tools. Yet research has largely passed over this medium: prior systems design bespoke memory representations and study retrieval over them, leaving the default's two working assumptions untested: that an agent can keep a growing store organized as memories accumulate, conflict, and go stale, and that this organization pays. We present the first systematic exploration of filesystem-based memory for LLM agents. We formalize the setting as three roles around one memory filesystem: a management agent integrates and organizes incoming content, a search agent answers queries with cited sources, and an execution agent supplies task trajectories that are distilled into skills, unifying declarative memory and skills in a single store. Across long-conversation benchmarks and embodied tasks, we vary memory shape (agent-organized hierarchy, verbatim dump, chunk retrieval), stream scale, tool harness (sandboxed shell, memory-tool-style functions, varied search tooling), and the strengths of the management and search agents, tracking answer quality, cost, and store health as memory grows. What organization reliably buys is search economy: organized stores roughly halve retrieval cost where material is large. Today's agents, however, fall short of the default's promise: in our growth study, organization erodes for all but the strongest management agent, and no agent we measure converts organization itself into better answers. And the model is not the only lever over a store's shape: changing the tool set alone reshapes the store as strongly as swapping the model. The study turns the filesystem default from an assumption into a design space for agent memory.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • University of California, San Diego(加利福尼亚大学圣迭戈分校)
  • University of California, Merced(加利福尼亚大学默塞德分校)
  • Adobe Research(奥多比研究院)
  • Texas A&M University(德克萨斯农工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑