arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

理解在早期完成:大语言模型的深度分工及其在无界上下文记忆中的应用

CoMem: Reusing Transformer Depth across Queries with Persistent Intermediate Residuals

Hanzuo Liu, Xuan Qi, Chunyu Liu, Haotian Zhong, Yulong Wang, Key, Rayying, Alex Lamb, Mingyu Gao

arXiv 2607.28263首次发表:更新:

发表机构

Tsinghua University; Tencent(清华大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出CoMem方法,利用大语言模型的深度分工构建无界上下文记忆,在RULER、LoCoMo等基准上性能优于全上下文KV-Direct,内存占用低、预填充速度快,证明长上下文记忆可沿层轴组织。

AI 中文摘要

Transformer的深度并非被均匀使用:低层和中层构建语义表征,而高层则日益将这些表征专门化以用于预测。我们将这种分工转化为CoMem(理解记忆),该方法仅通过中间层写入每个上下文块,检索固定数量的缓存残差状态,并在所得集合上重新计算以查询为条件的高层。在固定检索预算下,模型侧的读取计算和内存与存储上下文长度无关。我们在统一的无聊天模板协议下评估了持续微调的Qwen3-8B基础大语言模型,其主干被冻结;旗舰模型仅在普通PG19上训练了秩为32的自蒸馏LoRA,我们单独报告了无适配器分支的结果。CoMem在RULER上达到97.05,在LoCoMo上达到38.27,而全上下文KV-Direct在LoCoMo上仅为34.59;对话记忆优势在对话集群重采样和独立评判下依然存在。在额外的长上下文和长文档任务上的结果揭示了有界检索的益处及其窗口内压缩损耗。受控深度扫描显示,更深的缓存会降低每个查询的重新计算量,但会产生保真度损失,而自蒸馏可大幅修复该损失。在NVIDIA H20上128k的无适配器效率控制实验中,CoMem使用18.26 GB内存,而KV-Direct使用89.36 GB,预填充速度提升7.83倍。这些结果表明,长上下文记忆可沿层轴而非仅沿词元轴组织。

英文摘要

Repeated queries over shared documents repeatedly execute the same lower transformer layers. We introduce CoMem, which makes split depth j an explicit reusable-context axis: write one depth-j residual per token, select a bounded chunk set, and resume only layers [j:L). Among document-reuse systems we are aware of, CoMem jointly makes split depth a tunable serving axis and isolates it with a matched j=0 endpoint. On Qwen3-8B, j=12 reduces selected-pack Read from 931.9 to 664.4 ms (1.403x), with a 3.12-point RULER cost (95% CI [2.36, 3.93]); a continuous-prefix oracle recovers the full gap. The resulting depth axis quantifies a quality-latency-storage trade-off; a separate same-adapter, Write-inclusive pipeline is 2.74x faster. Equal-latency raw replay leads by 11.56 points with BM25, directly measuring an applicability boundary of prepaid depth rather than hiding it. CoMem stores 8 KiB/token versus 144 KiB/token for a protocol-aligned same-Qwen3 CacheBlend-style diagnostic; the cohorts and adaptation budgets are not matched. A context-position factorization identifies missing lower-layer document context as the dominant tested multikey error, and a 32-token overlap raises 92.5 to 98.5 without increasing persistent bytes or per-query Read. CoMem opens transformer depth as a measurable, tunable dimension for repeated-query long-context serving.

Comments32 pages, 2 figures. Published at the COLM 2026 Workshop on Efficient Reasoning

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑