arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PReCache:通过低秩预计算与中性重建实现多LoRA智能体的高效KV缓存共享

PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction

Hyesung Jeon, Hyeongju Ha, Jae-Joon Kim

arXiv 2609.34054首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PReCache通过低秩预计算和中性重建实现多LoRA智能体的KV缓存共享,无需训练,显著提升推理速度与吞吐量,同时保持高准确性。

AI 中文摘要

多LoRA智能体系统通过共享一个通用骨干模型实现高效的角色专业化。然而,每个智能体都会重复处理不断增长的共享轨迹,并构建自己的KV缓存,这在长时程任务中引入了大量的内存和计算冗余。现有的KV缓存共享方法减少了这种重复的预填充,但它们要么需要额外的训练或架构约束,要么保留了大量的模型计算。此外,直接的缓存重用会导致当前智能体依赖由前一个智能体的适配器生成的缓存状态,从而削弱了其自身LoRA编码的角色特定行为。我们提出了PReCache,一个无需训练的KV缓存共享框架,包含两个设计,即PreLRShared和ReBaseShared,它们共享使用预训练权重计算的基础缓存,并预计算一个紧凑的智能体特定低秩(LR)缓存。为了消除重复的预填充,PreLRShared在共享上下文首次被处理时预计算每个智能体的LR缓存,允许当前智能体使用自己的LR缓存,而无需重新处理先前智能体处理过的上下文。为了提高共享准确性,ReBaseShared从无适配器的隐藏状态重建共享的基础缓存,减少了由前一个智能体的适配化表示引起的剩余误差。为了最小化其重建成本,我们提出了两种推理方案,分别针对单流推理和并发服务,在每次智能体轮次之后或与其执行同时进行相同的重建。在多个模型和智能体基准上,与没有KV缓存共享的推理相比,PreLRShared实现了高达3.1倍的TTFT加速和2.3倍的每请求吞吐量提升。在评估的缓存共享方法中,ReBaseShared总体上最能保持准确性,相对于无缓存共享的推理,平均仅下降1.1个百分点。

英文摘要

Multi-LoRA agent systems enable efficient role specialization by sharing a common backbone model. However, each agent repeatedly processes the growing shared trajectory and constructs its own KV cache, introducing substantial memory and computation redundancy in long-horizon tasks. Existing KV cache sharing methods reduce this repeated prefill, but they either require additional training or architectural constraints or retain substantial model computation. Moreover, direct cache reuse causes the current agent to rely on cache states generated by the previous agent's adapter, weakening the role-specific behavior encoded by its own LoRA. We present PReCache, a training-free KV cache sharing framework with two designs, namely PreLRShared and ReBaseShared, that share the base cache computed using the pretrained weights and precompute a compact agent-specific low-rank (LR) cache. To remove repeated prefill, PreLRShared precomputes each agent's LR cache when the shared context is first processed, allowing the current agent to use its own LR cache without reprocessing context processed by previous agents. To improve sharing accuracy, ReBaseShared reconstructs the shared base cache from adapter-free hidden states, reducing the remaining error caused by the previous agent's adapted representation. To minimize its reconstruction cost, we propose two inference schemes tailored to single-stream inference and concurrent serving, performing the same reconstruction after each agent's turn or alongside its execution, respectively. Across multiple models and agent benchmarks, PreLRShared achieves up to a 3.1x TTFT speedup and a 2.3x improvement in per-request throughput over inference without KV cache sharing. ReBaseShared best preserves accuracy overall among the evaluated cache-sharing methods, with an average drop of only 1.1 points relative to inference without cache sharing.

Comments24 pages, 8 figures, 13 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑