arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17109cs.AIcs.CL

共享前缀KV缓存跨标准LoRA适配器复用:质量与服务权衡

Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs

  • AltSlate Labs LLP

机构由 AI 辅助整理,请以论文原文为准。

Dushyant Rajput

AI总结:

研究标准LoRA适配器复用共享骨干KV缓存的质量与服务权衡,发现全前缀复用质量损失小但内存节省未实现,热缓存TTFT收益随上下文增长。

AI中文摘要:

一种常见的小型模型部署方式是运行一个共享骨干网络,并配备多个针对相同上下文回答问题的LoRA专家。朴素地服务这些专家时,会为每个专家重新对共享上下文进行预填充。我们研究一个狭窄而实际的问题:对于已经训练好的标准LoRA适配器——而非为缓存兼容性重新训练的适配器——如果骨干网络的预填充KV缓存被计算一次并在各专家间复用,任务质量能保留多少,这又能带来多少服务成本的节省?在Qwen3-1.7B骨干网络上,使用两个适配器(在HotpotQA上进行抽取式问答,在GSM8K上进行算术推理),我们扫描了专家从复用基础缓存接管的分界点,并测量了成对的质量差异和服务成本。全前缀复用的预填充成本最低,在留出GSM8K上的质量差异很小(在160 token预算下Delta = -4.6 EM;在320 token下为-3.0;在第二个训练种子下为-0.8——所有这些都偏向原生方法,只有第一个排除零,且幅度不一致)。部分重计算未显示出明显优势。既未建立质量等价性,也未建立通用的分界点选择规则。我们还报告了一种闭式岭KV翻译器,它并未胜过直接复用,以及专家依赖性的对比,其置信区间均包含零。实测的服务收益是热缓存的首次令牌时间,它随上下文增长(在8K时约16倍);两分支的峰值内存仅降低12%,经检查,前缀从未在分支间物理共享——该实现复用了KV值但复制了其存储,因此未实现共享缓存的内存节省。

英文摘要:

A common small-model deployment runs one shared backbone with several LoRA specialists that answer over the same context. Serving them naively re-prefills that shared context once per specialist. We study a narrow, practical question: for already-trained standard LoRA adapters -- not adapters retrained for cache compatibility -- how much task quality is preserved if the backbone's prefill KV cache is computed once and reused across specialists, and what does that buy in serving cost? On a Qwen3-1.7B backbone with two adapters (extractive QA on HotpotQA, arithmetic reasoning on GSM8K), we sweep the boundary at which the specialist takes over from the reused base cache and measure paired quality differences and serving cost. Full-prefix reuse had the lowest prefill cost and a small quality difference on held-out GSM8K (Delta = -4.6 EM at a 160-token budget; -3.0 at 320 tokens; -0.8 under a second training seed -- all favoring native, only the first excluding zero, and the magnitude not consistent). Partial recomputation provided no demonstrated advantage. Neither quality equivalence nor a general boundary-selection rule is established. We also report a closed-form ridge KV translator that did not beat direct reuse, and specialist-dependence contrasts whose intervals all include zero. The measured serving benefit is warm-cache time-to-first-token, which grows with context (~16x at 8K); two-branch peak memory was only 12% lower and, on inspection, the prefix was never physically shared across branches -- this implementation reuses KV values but copies their storage, so shared-cache memory savings are not achieved.

↑