发表机构
UNSW Sydney(悉尼新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究验证了GLM-5.3-Flash在vLLM和LMCache下的混合状态缓存恢复修复,通过严格前缀查找对齐,将串行工作负载生成一致性从34/36提升至36/36,并显著降低首令牌延迟。
AI 中文摘要
外部缓存传输可能在混合语言模型从不一致状态恢复时成功。我们检查了完整的45层GLM-5.3-Flash模型,使用RedHatAI/GLM-5.3-Flash-NVFP4量化检查点,在四路张量并行下配合vLLM和LMCache。一次完全命中的恢复不匹配为完整提示恢复了状态,而调度器少计了一个令牌。我们通过严格前缀查找对齐恢复,并利用共享计算修正、匹配的检查点调度和固定的每秩内核配置建立了数值比较。在九长度串行工作负载中,与修改后的重计算控制的一致性从34/36提高到36/36次生成,每次生成包含64个令牌ID。一次单独的仪器化运行通过了记录的传输页、有效尾部和延迟保存检查。三个额外的合成模板在两次全新容器运行中通过了72对256令牌的续写。随后的串行性能研究在120个请求中保持了输出等价性;在测量的试验中,相对于修改后的冷重计算,CPU重载将首令牌时间减少了46-64%,总请求时间减少了1.9-7.0%。贡献是应用现有检查点对齐原则的实验验证的集成修复。证据仅限于一个模型版本和受控配置;它不确立通用确定性、任务质量等价性、并发服务收益或超出GPU内存的容量。
英文摘要
External cache transfers can succeed while a hybrid language model resumes from an inconsistent state. We examine the full 45-layer GLM-5.3-Flash model, using the RedHatAI/ GLM-5.3-Flash-NVFP4 quantized checkpoint with vLLM and LMCache under four-way tensor parallelism. A complete-hit recovery mismatch restored state for the full prompt while the scheduler credited one fewer token. We aligned recovery through strict-prefix lookup and established a numerical comparison using shared computation corrections, matched checkpoint scheduling, and fixed per-rank kernel configurations. In a nine-length serial workload, agreement with the modified recomputation control improved from 34/36 to 36/36 generations, each containing 64 token IDs. A separate instrumented run passed recorded transfer-page, effective-tail, and delayed-save checks. Three additional synthetic templates passed 72 paired 256-token continuations across two fresh-container runs. A subsequent serial performance study preserved output equality across 120 requests; among the measured trials, CPU reload reduced time to first token by 46-64% and total request time by 1.9-7.0% relative to modified cold recomputation. The contribution is an experimentally validated integration repair applying an existing checkpoint-alignment principle. The evidence is confined to one model revision and controlled configuration; it does not establish general determinism, task-quality equivalence, concurrent-serving gains, or capacity beyond GPU memory.
Comments8 pages, 3 figures, 7 tables