arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

现在正确,之后不足:审计上下文压缩中的更新充分性

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

Guangzhe Zhang

arXiv 2609.20045首次发表:更新:

发表机构

Independent AI Researcher(独立AI研究员)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过配对历史审计方法,揭示上下文压缩中记忆虽能正确回答当前查询但丢弃未来更新所需区分信息的问题,并提供了范围限定的评估框架与可复现的失败分析。

AI 中文摘要

一个记忆可以正确回答当前查询,同时丢弃后续更新所需的区分信息。我们通过配对历史审计来研究这一失败:两个历史具有相同的当前答案,接收共享的未来更新,并需要不同的后续答案。一项试点评估了跨六种合成机制、12种记忆条件、两次重复和两个模型后端的24对历史。一个确定性前沿选择器在DeepSeek上获得96/96的严格揭示准确率,在GLM上获得82/96;一个结构化写入器获得62次成功,1个未解决结果和56/96。配置的四结果联合对比具有有限样本识别区间[0.521, 0.542]和[0.292, 0.313],而非置信区间。一项记录级审计区分了保留状态充分性、响应传递和答案模式合规性,而不改变那些原始分数。它发现26个和25个格式良好但语义错误的结构化揭示记忆,而所有14个GLM前沿揭示失败在错误包装中包含正确值。墓碑移除在目标机制中产生16/16的精确重放失败。标识符重命名随后暴露了一个单独的缺陷:原始前沿晚期引用充分性从8/8下降到94/320个转换实例。我们提供并测试了一种标签等变修复,但它仅保留2/8个原始晚期引用答案:消除命名捷径并不能解决未知的未来相关性。这些结果支持一种范围限定的评估方法和可复现的失败分析,而非修复算法的普遍优越性。付费试点证据、回顾性诊断和新的离线测试分别报告;不声称独立的留出或自然任务验证。

英文摘要

A memory can answer a current query correctly while discarding distinctions required by a later update. We investigate this failure with a paired-history audit: two histories have the same current answer, receive a shared future update, and require different subsequent answers. A pilot evaluates 24 history pairs across six synthetic mechanisms, 12 memory conditions, two repeats, and two model backends. A deterministic frontier selector obtains strict reveal accuracy of 96/96 on DeepSeek and 82/96 on GLM; a structured writer obtains 62 successes with one unresolved outcome and 56/96. The configured four-outcome joint contrast has finite-sample identification intervals of [0.521, 0.542] and [0.292, 0.313], not confidence intervals. A record-level audit distinguishes retained-state adequacy, response delivery, and answer-schema compliance without changing those original scores. It finds 26 and 25 well-formed but semantically wrong structured reveal memories, while all 14 GLM frontier reveal failures contain correct values in the wrong wrapper. Tombstone removal produces 16/16 exact replay failures in the targeted mechanism. Identifier renaming then exposes a separate flaw: original frontier late-reference adequacy falls from 8/8 to 94/320 transformed instances. We provide and test a label-equivariant repair, but it preserves only 2/8 original late-reference answers: eliminating a naming shortcut does not solve unknown future relevance. These results support a scoped evaluation methodology and reproducible failure analysis, not general superiority of the repaired algorithm. Paid pilot evidence, retrospective diagnostics, and new offline tests are reported separately; no independent held-out or natural-task validation is claimed.

Comments20 pages, 9 tables, 2 figures. Code and reproducibility materials to be released separately

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑