有界状态恢复:将本地恢复能力与外部大语言模型状态解耦
Bounded-State Restoration: Decoupling Local Restore Capacity from External LLM State
浏览论文内容
中文总结 AI 辅助
提出有界状态恢复(BSR)方法,解耦本地恢复能力与外部LLM状态,在DeepSeek-V4-Flash上实现大外部状态下低本地RWS,缩短恢复TTFT。
中文摘要 AI 辅助
分层KV缓存系统可在GPU内存之外保留长上下文大语言模型(LLM)的执行状态,但保留能力并不决定使该状态重新可执行所需的本地内存。我们将这第二种资源隔离为恢复工作集(RWS):恢复期间生命周期重叠的峰值本地暂存状态。在固定上游LMCache全计划路径中,针对1.956、7.823和15.646 GiB/秩的状态,首次成功的全复用点对应2、8和16 GiB的L1层级,成功的L1峰值为1.956、7.824和15.648 GiB/秩。我们提出有界状态恢复(BSR),其将完整发现与本地驻留分离。BSR探测完整可复用前缀,无需在L1中实例化全部命中内容,随后通过最多W个块的可复用窗口安装已确认状态。在有界辅助状态下,峰值恢复能力为O(W),而总传输和安装工作仍为Θ(|S|)。由于可复用状态跨越异构分配器组和张量并行(TP)秩,BSR采用请求级提交规则:部分安装绝不会作为有效可复用前缀暴露;失败则使已公布前缀失效,并回退至较低有效层级或确定性重计算。在两个DGX Spark节点上运行TP=2的DeepSeek-V4-Flash时,无恢复的干净扫描将外部状态从1.956 GiB/秩扩展至31.277 GiB/秩,而在W=32时测得的L1 RWS始终为500.75 MiB/秩,实现了最大外部状态与活动暂存的63.959倍比率。第二次新的524K token运行重复了最大状态接受结果。对层级和秩非对称故障的评估显示,回退前要么完全复用,要么零外部复用。匹配的SSD优化将512K token恢复的首词生成时间(TTFT)从43.1秒缩短至17.6秒,且不改变RWS。
英文摘要
Hierarchical KV-cache systems can retain long-context LLM execution state beyond GPU memory, but retention capacity does not determine the local memory required to make that state executable again. We isolate this second resource as the restoration working set (RWS): the peak local staging state whose lifetimes overlap during restoration. In the pinned upstream LMCache whole-plan path, measured full-reuse points for 1.956, 7.823, and 15.646 GiB/rank states first succeed at 2, 8, and 16 GiB L1 rungs, with successful L1 peaks of 1.956, 7.824, and 15.648 GiB/rank. We introduce Bounded-State Restoration (BSR), which separates complete discovery from local residency. BSR probes the complete reusable prefix without materializing the whole hit in L1, then installs confirmed state through a reusable window of at most $W$ chunks. Under bounded auxiliary state, peak restoration capacity is $O(W)$ while total transfer and installation work remains $Θ(|S|)$. Because reusable state spans heterogeneous allocator groups and tensor-parallel ranks, BSR uses a request-level commit rule: partial installation is never exposed as a valid reusable prefix; failures invalidate the advertised prefix and fall back to a lower valid tier or deterministic recomputation. On DeepSeek-V4-Flash with TP=2 across two DGX Spark nodes, a clean no-resume sweep grows external state from 1.956 to 31.277 GiB/rank while measured L1 RWS remains exactly 500.75 MiB/rank at $W=32$, a 63.959x largest-state external-to-live-staging ratio. A second fresh 524K-token run repeats the largest-state acceptance result. Evaluated tier and rank-asymmetric failures expose either complete reuse or zero external reuse before fallback. A matched SSD optimization reduces 512K restore TTFT from 43.1 to 17.6 seconds without changing RWS.