arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SoloQ:扩散语言模型的无校准量化

SoloQ: Calibration-Free Quantization for Diffusion Language Models

Donghyun Lee, Arkapravo Ghosh, Varun Manjunath, Bumjoon Kyle Rhee, Hyunho Kook, Shiting Xiao, Youngeun Kim, Priyadarshini Panda

arXiv 2610.07121首次发表:更新:

发表机构

University of Southern California; Yale University; Korea University(南加州大学; 耶鲁大学; 高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SoloQ是一种无校准量化框架,通过旋转基映射实现数据无关量化,适用于扩散语言模型,在4位量化下保持精度并显著降低内存和加速推理。

AI 中文摘要

扩散大语言模型(dLLMs)通过基于双向扩散的令牌生成,已成为自回归语言模型的一种有前景的替代方案。然而,其不断增大的模型规模和高推理成本使得高效部署面临挑战:全序列去噪反复调用计算密集的前向传播,而块扩散模型还额外引入了内存密集的KV缓存。因此,低位宽权重激活量化具有吸引力,但现有的dLLM后训练量化方法依赖校准数据,尽管激活分布会随掩码状态和去噪步骤而变化。我们提出SoloQ,一种无校准量化框架,将权重和激活映射到具有可预测边际分布的归一化旋转基中,从而实现数据无关的量化。SoloQ结合了结构化K-RPBH旋转和轻量级重缩放校正,以实现无校准量化。其可预测的旋转后分布支持分布匹配码本和硬件原生NVFP4。对于块扩散模型,SoloQ进一步应用提交时KV缓存量化,以压缩持久状态而不干扰正在去噪的块。在全序列dLLMs(LLaDA和Dream)和块扩散dLLMs(Fast-dLLM v2和Nemotron-Labs-Diffusion)上,SoloQ在4位量化下保持准确性,并在知识和推理密集型基准上优于基于校准的基线。使用NVFP4,SoloQ将峰值内存减少高达2.61倍,并将端到端推理加速高达2.24倍。

英文摘要

Diffusion large language models dLLMs) have emerged as a promising alternative to autoregressive language models through bidirectional diffusion-based token generation. However, their growing model sizes and high inference costs make efficient deployment challenging: full-sequence denoising repeatedly invokes compute-intensive forward passes, while block-diffusion models additionally introduce a memory-intensive KV-cache. Low-bit weight-activation quantization is therefore attractive, yet existing dLLM post-training quantization methods rely on calibration data despite activation distributions shifting across masking states and denoising steps. We present SoloQ, a calibration-free quantization framework that maps weights and activations into a normalized rotated basis with a predictable marginal distribution, enabling data-independent quantization. SoloQ combines a structured K-RPBH rotation with a lightweight rescaling correction for calibration-free quantization. Its predictable post-rotation distribution supports both distribution-matched codebooks and hardware-native NVFP4. For block-diffusion models, SoloQ further applies commit-time KV-cache quantization to compress persistent states without perturbing the actively denoised block. Across full-sequence dLLMs (LLaDA and Dream) and block-diffusion dLLMs(Fast-dLLM v2 and Nemotron-Labs-Diffusion), SoloQ retains accuracy under 4-bit quantization and outperforms calibration-based baselines on knowledge- and reasoning-intensive benchmarks. With NVFP4, SoloQ reduces peak memory by up to 2.61X and accelerates end-to-end inference by up to 2.24X.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑