发表机构
Cornell University; Arizona State University(康奈尔大学; 亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LatentPress将上下文压缩为连续记忆令牌,仅训练小型适配器,在多个长上下文任务中压缩后性能优于文本摘要和OCR,且读写速度更快,证实其作为机器面向上下文接口的实用性。
AI 中文摘要
压缩上下文通常以人类可读的文本或渲染图像形式存在,即便其使用者是语言模型,也必须对这些内容进行解码。本文介绍LatentPress,它将对话历史和长文档写入第三种表示形式:连续记忆令牌,冻结的解码器可通过其输入嵌入接口直接读取这些令牌,推理过程中无需文本重构。一个小型的与读取器匹配的写入器可实现4-16倍的压缩率,且仅训练适配器(参数规模为420万至2620万,约为解码器参数的0.1%)。在LongMemEval任务中,LatentPress在7.70倍压缩率下达到0.504的准确率,而未压缩证据的准确率为0.490,其表现优于文本摘要(准确率0.184)和基于OCR的压缩(准确率0.426至0.312)。在LongBench-QA任务中,域内写入器在4-8倍压缩率下的表现与原始上下文读取相当或更优,而16倍压缩时的表现略逊于原始上下文。每次对话的写入耗时为43毫秒,比文本摘要或OCR重构快约一个数量级,读取速度则比原始上下文或缓存OCR快5-9倍。我们在两种迁移设置下验证了该接口:从UltraChat到LongMemEval记忆问答的零样本迁移,以及从LongMemEval衍生的问答到未见过的LongBench文档域的迁移,证实了直接软令牌作为超越文本和视觉的机器面向上下文接口的实用性。实验的实现可在此URL获取。
英文摘要
Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even when its consumer is a language model. We introduce LatentPress, which writes conversational histories and long documents into a third representation: continuous memory tokens that a frozen decoder reads directly through its input-embedding interface, with no text reconstruction at inference. A small reader-matched writer compresses $4$-$16\times$ while training only an adapter (4.2M-26.2M parameters, $\sim\!0.1\%$ of the decoder). On LongMemEval, LatentPress reaches $0.504$ accuracy at $7.70\times$ compression versus $0.490$ for uncompressed evidence, outperforming text summaries (0.184) and OCR-based compression (0.426 to 0.312). On LongBench-QA, in-domain writers match or exceed raw-context reading at $4$-$8\times$ compression, while $16\times$ trails raw. Writing takes 43ms per conversation, roughly an order of magnitude faster than text summarization or OCR reconstruction, and reading is $5$-$9\times$ faster than raw context or cached OCR. We validate the interface under two transfer settings, zero-shot from UltraChat to LongMemEval memory QA and from LongMemEval-derived QA to unseen LongBench document domains, establishing direct soft tokens as a practical machine-facing context interface beyond text and vision. The implementation of the experiments could be found at: https://github.com/HJSang/LatentPress .