AI 中文总结
本研究通过EverMemBench受控实验发现,增加上下文宽度可显著提升LLM长期记忆检索准确率(最高17.98个百分点),而增加深度无单调收益,建议采用大型连贯上下文块替代深层记忆结构。
AI 中文摘要
长期对话记忆正成为现代LLM系统不可或缺的组成部分。已有架构按主题和事件对记录进行分组,构建层次结构和图结构,并通过因果和时间关系连接事实。我们通过实验研究了两个记忆参数之间的交互作用:结构深度和提供给答案模型的上下文宽度。使用EverMemBench,我们评估了深度D1-D4、1,024/2,048/4,096个token的核心预算,以及扩展到完整档案的额外Production和Oracle条件。将宽度从1K增加到4K,准确率提升了10.11-17.98个百分点,而增加深度并未带来单调增益。超过8-16K后,Production性能达到平台期,而每个正确答案的token数持续增加;Oracle在68-71K token的完整档案上保持质量。这些结果促使我们进一步研究大型连贯上下文块,而非逐步加深的记忆结构。
英文摘要
Long-term conversational memory is becoming an integral component of modern LLM systems. Proposed architectures group records by topics and events, construct hierarchies and graphs, and connect facts through causal and temporal relations. We experimentally study the interaction between two memory parameters: structural depth and the width of context supplied to the answer model. Using EverMemBench, we evaluate depths D1-D4, core budgets of 1,024/2,048/4,096 tokens, and additional Production and Oracle conditions up to the full archive. Increasing width from 1K to 4K improves Accuracy by 10.11-17.98 percentage points, whereas increasing depth provides no monotonic gain. Beyond 8-16K, Production performance reaches a plateau while tokens per correct answer continue to increase; Oracle preserves quality on full archives of 68-71K tokens. These results motivate further investigation of large, coherent context blocks instead of progressively deeper memory structures.
Comments4 pages, 1 figure. Accepted at the PALM Workshop at NeurIPS 2026