arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11775cs.AIcs.CL

休眠智能体:基于主旨的上下文压缩会丢失什么以及为什么

The Sleeping Agent: What Gist-Based Context Compression Loses and Why

Nicholas E. Kyrkewood

AI总结:

本研究以SWC为工具,发现主旨压缩对多跳推理等任务表现优于截断,但会损害时间相关问题的性能,通过修改提示可提升时间问题的评判准确率。

AI中文摘要:

基于主旨的上下文压缩——将较早的对话历史总结为紧凑表示——是长程语言模型智能体的常用方法,但人们对其对不同类型记忆检索的影响知之甚少。我们采用Salience-Weighted Consolidation(SWC,一种受睡眠记忆整合启发的生物驱动压缩框架)作为诊断探针,研究主旨压缩何时有益、何时有害。SWC按显著性对对话历史评分,将其划分为优先级层级,并对中等优先级内容应用结构化主旨抽象。在所有10个LoCoMo对话(共1935个匹配的纯文本问题,排除第5类(对抗性)问题后,主要聚合分析使用1501个问题)上以温度0评估4种条件,我们发现一致的任务类型交互作用:主旨压缩在多跳推理和单跳事实问题上的表现显著优于截断,但在时间相关问题上,压缩条件的表现仍明显更差,在同时评估压缩条件和完整上下文的对话中,压缩条件的得分远低于完整上下文基准。我们将这种失败归因于特定机制:主旨抽象提示保留了关系和事件结构,却丢弃了日期和时间。对所有10个对话的保留分析证实了该机制:通过一句提示修改,时间表达的保留率从3.05%提高到62.39%(约20倍),而命名实体和事件的保留率几乎没有变化(分别为1.02倍和1.11倍),表明该修复是一种精准工具。该提示修改使匹配集中第2类(时间相关)问题的评判准确率提高了0.314[0.254, 0.375]。代码和结果:this https URL。

英文摘要:

Gist-based context compression---summarising older conversation history into compact representations---is a common approach in long-horizon language model agents, yet its effect on different types of memory retrieval is poorly understood. We use Salience-Weighted Consolidation (SWC), a biologically-inspired compression framework motivated by sleep-based memory consolidation, as a diagnostic probe to study when gist compression helps and when it hurts. SWC scores conversation history by salience, partitions it into priority tiers, and applies structured gist abstraction to mid-priority content. Evaluating four conditions on all ten LoCoMo conversations---1,935 matched text-only questions in total, 1,501 used in the primary aggregate after excluding Category 5 (adversarial) questions---at temperature 0, we find a consistent task-type interaction: gist compression substantially outperforms truncation on multi-hop reasoning and single-hop factual questions, but temporal questions remain substantially harder under compression, with compressed conditions scoring well below the full-context reference on the conversations where both are evaluated. We trace this failure to a specific mechanism: the gist abstraction prompt preserves relational and event structure while discarding dates and times. A preservation analysis across all ten conversations confirms the mechanism: an approximately 20-fold increase in temporal expression preservation (3.05% to 62.39%) with a one-sentence prompt modification, while named entity and event preservation rates barely change (x1.02 and x1.11), demonstrating that the fix is a precision instrument. The prompt modification recovers +0.314 [0.254, 0.375] judge accuracy on category-2 (temporal) questions in the matched set. Code and results: https://github.com/kyrkewood/sleeping-agent.

补充信息

↑