发表机构
Metask Lab(Metask实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
WitCert为KV缓存量化提供可靠运行时度量与门控,可实时观测压缩风险,恢复FP8量化质量并提升INT8缓存的KV令牌服务量,核心定理已在Lean 4中验证。
AI 中文摘要
KV缓存量化目前仅通过离线基准平均值验证,部署的系统无法判断压缩是否会损害当前正在处理的请求。我们提供了一种可证明可靠的运行时度量工具,即“KV量化的DTrace”:它针对每(层、头、步)给出精确注意力与压缩注意力之间总变差的上界。该度量工具分为两层:确定性的带范数见证界,对任何保留缓存的黑盒量化器及任何查询(自适应安全、最坏情况柯西-施瓦茨不等式加旋转位置编码带酉性)均可靠;以及针对受控减法抖动INT8量化器的更紧概率证书,适用于显式请求级失败预算(针对非自适应查询表述,核心定理已在Lean 4中通过机器验证)。三项结果:可观测性方面,该度量工具通过环境保护补丁接入SGLang,任何注册为张量函数的方案均可在实时服务中被测量;修复方面,由度量驱动的门控机制,在见证饱和时进行风险排序、在证书有效时进行门控,经验证可在基准规模上恢复质量下限,例如原始转换FP8在困难RULER任务上从22.8回升至79.7,配对测试显示其与未压缩结果的差异被限制在[+0.0, +0.8];分析方面,激进方案依赖跨层误差抵消而非每步保真度,在28层扫描中,无任何单层污染单独导致性能损失(0/28),且经过验证的INT8缓存可在SGLang中以相同内存服务1.88倍的KV令牌。
英文摘要
KV-cache quantization is validated today by offline benchmark averages; a deployed system cannot tell whether compression is damaging the request it is serving right now. We give it a provably sound runtime meter -- a "DTrace for KV quantization": a per-(layer, head, step) upper bound on the total variation between exact and compressed attention. The meter has two tiers: a deterministic band-norm-witness bound, sound for any cache-preserving black-box quantizer and for any query (adaptive-safe, worst-case Cauchy--Schwarz plus RoPE band-unitarity), and a tighter probabilistic certificate for a controlled subtractively-dithered INT8 quantizer under an explicit request-level failure budget (stated for non-adaptive queries; core theorems machine-checked in Lean 4). Three results. Observability: the meter enters SGLang through an env-guarded patch, and any scheme registered as one tensor function is measured in live serving. Repair: meter-driven gating -- risk-ranked where the witness is saturated, certified where it is informative -- empirically restores the quality floor at benchmark scale, e.g. raw-cast fp8 from 22.8 back to 79.7 on hard RULER tasks with the difference from uncompressed bounded at $[+0.0,+0.8]$ by a paired test. Analysis: aggressive schemes survive on cross-layer error cancellation, not per-step fidelity -- in a 28-layer sweep, no single layer's pollution alone loses anything (0/28) -- and the certified int8 cache serves $1.88\times$ more KV tokens at the same memory in SGLang. All artifacts, guards, and the Lean development are released at https://github.com/metask-ai/witcert-kv-certificates; every number regenerates from the shipped artifacts by one command.
Comments39 pages, 7 figures. Code, artifacts, and Lean proofs: https://github.com/metask-ai/witcert-kv-certificates