发表机构
University of Southern California; University of California, Irvine; Intel Labs(南加州大学; 加利福尼亚大学欧文分校; 英特尔实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RaReCache是一种基于秩分歧选择性重计算的框架,可让小型模型预填缓存后,大型模型仅重计算关键token即可准确解码,大幅降低跨模型KV缓存复用的预填延迟,提升服务吞吐量。
AI 中文摘要
跨模型KV缓存复用仍是现代大语言模型(LLM)服务中的关键挑战。编码智能体与多模型系统日益在模型间路由共享上下文:用户可能在会话中途切换模型,或级联系统会对困难查询进行升级。由于KV缓存包含模型特定的表示,每次切换通常会迫使接收模型从头预填整个上下文。近期研究表明,闭式线性映射可在同一系列的模型间转换KV缓存,但随着模型规模差距扩大,转换精度会下降。本文中,我们证实这些转换失败集中在一小部分信息密集型token上。为弥合该差距,我们提出RaReCache,这一框架可让大型目标模型通过选择性重计算,从更小源模型预填的缓存中准确解码。RaReCache采用一种新颖的秩分歧指标识别这些关键位置,通过校准数据弱支持的输出方向中映射KV的能量对每个token打分。在两个模型系列和五个基准测试中,当参数差距达23倍(Qwen3-0.6B至14B)时,仅重计算30%的位置即可保留95%-99%的目标精度;当参数差距为8.8倍(Llama3-8B至70B)时,重计算40%的位置可保留96.5%的目标精度。RaReCache大幅降低了对源模型规模的敏感性,实现了最高3.04倍的预填加速。对于在线服务,在单GPU上其处理的请求吞吐量是目标预填的1.8倍,且在目标饱和负载下,通过30%的重计算预算,将首token时间(TTFT)的中位数和第99百分位数分别降低了5.0倍和6.4倍。RaReCache建立了一种高效的服务范式,其中小型模型代表大型目标进行预填,大型模型仅重计算关键token,大幅降低了预填延迟。
英文摘要
Cross-model KV-cache reuse remains a key challenge in modern LLM serving. Coding agents and multi-model systems increasingly route a shared context across models: a user may switch models mid-session, or a cascade may escalate a difficult query. Because KV caches contain model-specific representations, each switch typically forces the receiving model to prefill the entire context from scratch. Recent work shows that closed-form linear maps can translate KV caches between models in the same family, but transfer accuracy degrades as the model-size gap widens. In this paper, we establish that these transfer failures are concentrated in a small subset of information-dense tokens. To bridge this gap, we introduce RaReCache, a framework that enables a large target model to decode accurately from a cache prefilled by a much smaller source via selective recomputation. RaReCache identifies these critical positions using a novel rank disagreement metric, scoring each token by the energy of its mapped KV in output directions weakly supported by the calibration data. Across two model families and five benchmarks, on a 23x parameter gap (Qwen3-0.6B to 14B) recomputing just 30% of positions retains 95-99% of the target accuracy, whereas on a 8.8x gap (Llama3-8B to 70B), recomputing 40% retains 96.5% of the target accuracy. RaReCache largely removes sensitivity to source-model size, and achieves up to a 3.04x prefill speedup. For online serving, it handles 1.8x the request throughput of target prefill on a single GPU, and at the target's saturation load, reduces median and 99th-percentile time-to-first-token (TTFT) by 5.0x and 6.4x respectively, with a 30% recompute budget. RaReCache establishes an efficient serving paradigm where small models prefill on behalf of massive targets, enabling large models to recompute only critical tokens, drastically reducing prefill latency.