arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

保留还是合并?语言智能体记忆中依赖预算的算子选择

Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory

Qingcan Kang, Mingyang Liu, Shixiong Kai, Kaichao Liang, Zhentao Tang, Yuqi Cui, Tao Zhong, Mingxuan Yuan

arXiv 2607.17545首次发表:更新:

AI 中文总结

研究语言智能体记忆中预算依赖的算子选择问题,核心方法是将算子效用分解,用离线抽象安全机制实现,主要贡献是通过实验揭示在不同预算下保留和合并策略的适用情况及算子效果差异。

AI 中文摘要

语言智能体在交互过程中依赖记忆。然而,大语言模型有限的上下文窗口及其推理成本限制了一次可使用的记忆量。现有系统主要采用两种策略:记忆保留和记忆合并。保留策略保留原始记录和确切细节,但在预算紧张时相关证据可能无法全部保留;合并策略压缩并组合记录,提高每个token的覆盖范围,但有丢失关键查询细节的风险。两种策略都并非普遍适用。这引发了两个核心问题:何时合并应取代保留,以及应选择哪种算子(合并、摘要或重写)?我们通过将每个算子的效用分解为对保留策略遗漏证据的覆盖效果和对已适配原始证据的有符号替换效果来形式化这一决策。它们的平衡解释了为何首选行动会随相对预算压力而变化。我们使用离线抽象安全(OAS)实现了这一机制,这是一种轻量级学习者,通过留出的危害校准从预生成特征中估计行动效用。公开的LongMemEval和LoCoMo基准显示出相同的预算依赖模式。在LongMemEval上,在预算紧张时合并策略可将绝对准确率提高多达48%,而在预算宽松时保留策略更可取;LoCoMo在较小预算下也呈现了这种交叉情况,与其较短的证据一致。在两个数据集上,当需要压缩时,跨笔记抽象和合并通常优于局部重写。

英文摘要

Language agents depend on memory across interactions. However, the limited context windows of large language models (LLMs) and their inference costs constrain how much memory can be used at once. Existing systems mainly follow two strategies: memory retention and memory consolidation. Retention keeps raw records and preserves exact details, but relevant evidence may not fit under a tight budget; consolidation compresses and combines records, improving coverage per token but risking the loss of query-critical details. Neither strategy is universally preferable. This raises two central questions: when should consolidation replace retention, and which operator -- Merge, Abstract, or Rewrite -- should be selected? We formalize this decision by decomposing each operator's utility into a coverage effect on evidence omitted by retention and a signed replacement effect on raw evidence that already fits. Their balance explains why the preferred action changes with relative budget pressure. We implement this mechanism with Offline Abstraction-Safety (OAS), a lightweight learner that estimates action utilities from pre-generation features with held-out harm calibration. The public LongMemEval and LoCoMo benchmarks show the same budget-dependent pattern. On LongMemEval, consolidation improves absolute accuracy by up to 48% under tight budgets, whereas retention is preferable under loose budgets; LoCoMo replicates this crossover at a smaller budget, consistent with its shorter evidence. On both datasets, cross-note abstraction and merging generally outperform local rewriting when compression is necessary.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑