Zero-Mem:面向大语言模型智能体的零令牌内存操作
Zero-Mem: Zero-Token Memory Operations for LLM Agents
浏览论文内容
中文总结 AI 辅助
Zero-Mem提出零令牌内存操作,以原始交互轨迹为记录来源,通过实体-上下文图与时间层次结构组织信息,仅在最终问答时调用LLM,在长记忆问答基准中实现性能与效率的平衡,降低内存操作时间成本57.6%。
中文摘要 AI 辅助
大语言模型(LLM)智能体需要内存以在长交互中保持一致性,然而许多系统会使用额外的LLM调用来操作该内存。生成中间记录并协调其检索会产生持续的令牌和时间成本,而被遗漏或合并的细节可能会模糊原始证据。我们提出疑问:结构化内存访问是否完全需要生成?Zero-Mem引入了“零令牌内存操作”:除最终问答外的所有步骤均不调用LLM,也不消耗LLM的输入或输出令牌;编码器计算单独核算。Zero-Mem将原始交互轨迹作为记录来源,以两种互补方式组织这些轨迹:实体-上下文图揭示跨交互的连接,时间层次结构保留会话局部性和会话状态。对于每个查询,Zero-Mem权衡两种视图,从两者中检索,并遵循其结构恢复支持关系或周围上下文。确定性校准首先丢弃冲突证据,然后使读者的答案基于检索到的轨迹。仅最终问答(final-QA)读者会调用LLM。在长记忆和长上下文问答基准测试中,Zero-Mem在消除内存操作中的LLM调用和LLM令牌消耗的同时,取得了具有竞争力的性能。使用相同的final-QA读者和上下文预算,它相对于最快的对比基线,将内存操作的时间成本降低了57.6%。消融实验验证了两种视图及其依赖查询的协调的贡献。总体而言,结果表明,结构化智能体内存无需生成过去的中间表示。同行评审后,代码和实现细节将在该https URL处提供。
英文摘要
LLM agents need memory to act consistently over long interactions, yet many systems use additional LLM calls to operate that memory. Generating intermediate records and mediating their retrieval adds recurring token and time costs, while omitted or merged details can obscure the original evidence. We ask whether structured memory access requires generation at all. Zero-Mem introduces \emph{zero-token memory operations}: no step outside final question answering invokes an LLM or consumes LLM input or output tokens; encoder computation is accounted for separately. Zero-Mem preserves original interaction traces as its source of record. It organizes the traces in two complementary ways. An entity--context graph exposes connections across interactions, while a temporal hierarchy preserves conversational locality and session state. For each query, Zero-Mem weighs the two views, retrieves from both, and follows their structure to recover supporting relations or surrounding context. Deterministic calibration first discards conflicting evidence and then keeps the reader's answer grounded in the retrieved traces. Only the final-QA reader invokes an LLM. Across long-memory and long-context question-answering benchmarks, Zero-Mem achieves competitive performance while eliminating LLM calls and LLM-token consumption from memory operations. With the same final-QA reader and context budget, it reduces memory-operation time cost by 57.6\% relative to the fastest compared baseline. Ablations support the contribution of the two views and their query-dependent coordination. Overall, the results show that structured agent memory need not generate an intermediate representation of the past. After peer review, the code and implementation details will be available at \textcolor{blue}{https://github.com/TheMoon0815/Zero-mem}.