发表机构
Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EncBank通过复用预训练LLM下层编码输出并紧凑存储,实现跨查询高效推理,在保持任务性能的同时显著节省存储和加速预填充。
AI 中文摘要
对共享文档的重复查询会导致冗余编码,而缓存模型状态则带来持续的存储成本。基于CoMem的中间状态接口,EncBank将预训练LLM的下层视为可复用的文档编码器,并紧凑地存储其输出以供适配的上层阅读器使用。一种自蒸馏的后缀适配器在各骨干网络内的不同存储精度间共享,无需针对量化进行重新训练。在三个Qwen骨干网络(涵盖不同规模及全注意力和混合架构)的五个基准套件上,4位存储使每个报告的基准聚合值保持在原生精度EncBank的一个分数点以内。在固定的Qwen3-8B工作负载中,它保留了原生精度持久GPU存储的28.1%。独立的原生精度对照实验显示,与相同证据、相同适配器的文本重放相比,选定包预填充速度提升1.40倍,但RULER准确率代价为3.12分。原生精度的Qwen3.8-27B配置还通过了Terminal-Bench 2.1的89项任务中的70项。因此,EncBank将可复用计算与紧凑内存相结合,而任务保真度和端到端收益仍取决于工作负载、准备成本和复用频率。
英文摘要
Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-precision persistent GPU store. Separate native-precision controls yield a 1.40x selected-pack prefill speedup over same-evidence, same-adapter text replay, at a 3.12-point RULER accuracy cost. A native-precision Qwen3.8-27B configuration also passes 70 of 89 Terminal-Bench 2.1 tasks. EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.
Comments17 pages, 3 figures, 7 tables