arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26333cs.LG

分离量化:专门化LLM预填充与解码阶段

Disaggregated Quantization: Specializing LLM Prefill and Decode

Andrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi, Tijmen Blankevoort, Dan Alistarh

首次发表
浏览论文内容

中文总结 AI 辅助

提出分离量化(DQ)方法,为LLM预填充和解码阶段分别定制计算格式、权重与存储,在Qwen和Gemma模型上提升解码准确性、加速提示处理,并通过SSD卸载实现1.78倍TTFT加速。

中文摘要 AI 辅助

预填充(prefill)和解码(decode)阶段对量化方法有不同的需求:低精度算术可加速提示处理,而紧凑的权重可减少生成过程中的内存流量。我们提出“分离量化”(disaggregated quantization, DQ),该方法将计算格式、权重和存储布局分别专门化以适应这两个阶段。在Qwen 3和Gemma 3上,仅针对解码阶段移除激活量化,可在不增加推理成本的情况下提升解码密集型任务的准确性。训练独立的计算原生预填充权重,相比仅权重的推理,可加速提示处理,同时在2-3比特解码下,其准确性在解码密集型和预填充密集型任务上均达到或超过基线。利用已发布的Qwen3.8-27B GGUF解码器,训练一个NVFP4预填充器,在不修改解码检查点的情况下,将1比特准确性在MMLU-Pro上提升32.5个百分点,在MMMU-Pro上提升35.3个百分点。为在单设备上容纳额外的检查点,卸载式分离预填充(offloaded disaggregated prefill, ODP)从SSD流式加载权重,并将加载时间分摊到提示长度上。在相同的27B模型上,ODP在8K提示长度下,相比仅权重基线,实现了1.78倍的首次令牌时间(time-to-first-token)加速(见http URL)。我们在vLLM中评估了分离服务下的准确性,并通过后训练量化在高达2.8T参数的模型上进一步验证了共享权重格式分离的有效性。

英文摘要

Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases. On Qwen 3 and Gemma 3, removing activation quantization specifically on decode improves accuracy on decode-heavy tasks without increasing inference cost. Training separate compute-native prefill weights accelerates prompt processing relative to weight-only inference while matching or exceeding its accuracy at 2-3-bit decode on both decode-heavy and prefill-heavy tasks. With released Qwen3.8-27B GGUF decoders, training an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without modifying the decode checkpoint. To accommodate the additional checkpoint on a single device, offloaded disaggregated prefill (ODP) streams its weights from SSD, amortizing loading over prompt length. On the same 27B model, ODP delivers a 1.78x time-to-first-token speedup over the weight-only baseline at 8K prompt length in llama.cpp. We evaluate accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters.

发表机构

  • NVIDIA(英伟达)
  • ISTA(奥地利科学技术学院)

机构由 AI 辅助整理,请以论文原文为准。

↑