MEMO:面向长时程LLM智能体的多模态证据记忆组织
MEMO: Multimodal Evidence Memory Organization for Long-Horizon LLM Agents
浏览论文内容
中文总结 AI 辅助
针对长时程LLM智能体记忆组织难题,提出多模态证据记忆组织方法MEMO,通过证据提取、模态分配和布局选择,以更少令牌提升下游任务性能。
中文摘要 AI 辅助
长时间运行的LLM智能体依赖外部记忆来存储和重用超出单个上下文窗口的信息,然而,交互轨迹的持续积累与有限的上下文容量之间存在根本性矛盾。因此,智能体记忆的关键挑战不仅在于检索相关记录,还在于在给定预算下选择必要的证据,并以适当的模态进行组织。现有的记忆读出方法主要采用文本或视觉形式。文本能保持高保真度,但其线性令牌表示使得不同重要性的内容以几乎均匀的单位成本争夺有限的上下文。视觉读出将文本渲染为类似文档的图像,可利用二维布局来展现结构并强调关键信息,但在渲染和压缩过程中可能丢失细粒度细节。为解决这一问题,我们提出了MEMO,一种面向LLM智能体的多模态证据记忆组织方法。MEMO首先使用训练好的证据提取器选择相关记忆块,并形成带有来源信息和呈现要求的证据单元。一个经过训练、以查询为条件的记忆管理器将每个单元分配给文本、视觉或双通道载体,并选择与证据结构匹配的布局。随后,一个确定性的记忆构建模块生成文本包和视觉页面。记忆管理器通过离线阅读器的反馈进行训练,该阅读器衡量引导式记忆计划的效用,从而使保留和呈现决策与下游使用保持一致。我们在四个基准测试(HotpotQA、2WikiMultiHopQA、LoCoMo和ALFWorld)上,使用多种阅读器后端对MEMO进行了评估。结果表明,MEMO能以更少的记忆令牌更高效地呈现记忆,提升下游任务性能,并在受限预算下构建更有效的工作记忆。
英文摘要
Long-running LLM agents rely on external memory to store and reuse information beyond a single context window, yet there is a fundamental tension between the continuous accumulation of interaction trajectories and the limited context capacity. The key challenge in agent memory is therefore not only to retrieve relevant records, but also to select necessary evidence under a given budget and organize it in an appropriate modality. Existing memory readout methods mainly use textual or visual forms. Text preserves high fidelity, but its linear token representation makes contents with different importance compete for the limited context at nearly uniform unit cost. Visual readout renders text into document-like images, which can use two-dimensional layouts to expose structure and emphasize key information, but it may lose fine-grained details during rendering and compression. To address this issue, we propose MEMO, a multimodal evidence memory organization method for LLM agents. MEMO first uses a trained evidence extractor to select relevant memory blocks and form evidence units with source information and presentation requirements. A trained query-conditioned memory manager assigns each unit to a textual, visual, or dual-channel carrier and selects a layout that matches the evidence structure. A deterministic memory construction module then generates the textual package and visual pages. The memory manager is trained with feedback from an offline reader that measures the utility of the guided memory plan, so that retention and presentation decisions align with downstream usage. We evaluate MEMO on four benchmarks, HotpotQA, 2WikiMultiHopQA, LoCoMo, and ALFWorld, with multiple reader backends. The results show that MEMO presents memory more efficiently with fewer memory tokens, improves downstream task performance, and builds more effective working memory under constrained budgets.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Tsinghua University(清华大学)
- Institute of Automation of the Chinese Academy of Sciences(中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。