AI 中文总结
SemPIC通过训练启用LoRA的Writer编译文档KV,结合KV梯度检查点,提升了长上下文场景下的KV缓存性能,平均微F1值达0.60,接近全重计算效果。
AI 中文摘要
长上下文检索与智能体工作负载会在指令、历史及文档顺序变化时重复复用相同文档。前缀缓存无法利用这种复用,而位置无关缓存(PIC)因独立编译的键值(KV)状态缺乏其将被使用的未来上下文,仍不可靠。我们的诊断显示,学习到的边界条件基线大幅降低了可复用块边界附近的注意力偏差,但仍存在内部和任务级残差,这促使我们对文档表示本身进行适配。我们提出SemPIC,它训练一个启用LoRA的Writer,通过行为蒸馏编译原生的每层文档KV,同时保留预训练解码器作为不变的Reader。适配仅局限于离线缓存构建,保留标准KV接口和缓存命中解码路径。我们还引入KV梯度检查点,可降低峰值训练内存且不会切断通过缓存KV的梯度。在三个模型和四个任务上,SemPIC将KV Packet的平均微F1值从0.53提升至0.60,接近全重计算的0.62。
英文摘要
Long-context retrieval and agentic workloads repeatedly reuse the same documents under changing instructions, histories, and document orders. Prefix caching cannot exploit this reuse, while position-independent caching (PIC) remains unreliable because independently compiled KV states lack the future context in which they will be consumed. Our diagnostics show that a learned boundary-conditioned baseline sharply reduces attention deviation near reusable-block boundaries but leaves interior and task-level residuals, motivating adaptation of the document representation itself. We present \emph{SemPIC}, which trains a LoRA-enabled Writer to compile native per-layer document KVs through behavioral distillation while retaining the pretrained decoder as an unchanged Reader. Adaptation is confined to offline cache construction, preserving the standard KV interface and cache-hit decoding path. We further introduce KV Gradient Checkpointing, which reduces peak training memory without severing gradients through cached KVs. Across three models and four tasks, SemPIC raises mean micro-F1 over KV Packet from 0.53 to 0.60, approaching Full Recompute at 0.62. Code: https://github.com/jn12-29/SemPIC