arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为何门控DeltaNet在4位量化中仍能存活:混合27B大语言模型(LLM)循环部分的NVFP4 W4A4方案

Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM

Sergii Kozyrev, Davyd Maiboroda

arXiv 2609.04098首次发表:更新:

AI 中文总结

本研究针对混合27B LLM的循环部分,提出NVFP4 W4A4全线性层量化方案,通过机制分析验证其性能与BF16相当,且体积更小、预填充速度更快,明确了循环部分易量化的原因

AI 中文摘要

混合大语言模型(LLM)将softmax注意力与线性注意力层(如门控DeltaNet(GDN))结合,其循环状态以固定大小汇总上下文。早期社区对Qwen3.8-27B(含48个GDN层、16个注意力层)的4位量化操作,基于“循环误差会在长上下文中累积”的直觉,将GDN模块(尤其是其衰减门和写入强度门)保留在8位或16位精度。我们通过构建Minima方案验证该直觉:对所有496个线性层(含GDN)采用NVFP4 W4A4量化。在4K/32K困惑度、MMLU-Pro、GSM8K、AIME'25、GPQA-Diamond、LiveCodeBench及64K长度的RULER检索任务中,Minima的性能与BF16精度的模型处于种子噪声范围内(5项任务平均差距为-0.52),同时是我们对比方案中体积最小(17.5 GiB)、预填充速度最快(提升14%-19%)的方案,且其32K困惑度差距随位置缩小。四部分机制研究解释了原因:(i)NVFP4的16元素块缩放将残差流的极端异常值局部化,均衡了不同层角色的激活误差;(ii)原本被认为脆弱的门投影敏感度最低——softplus/指数和sigmoid参数化将约11%的通用矩阵乘(GEMM)误差压缩为约2%的输出误差;(iii)delta规则循环在32K token上将注入噪声维持在平坦平台,并在数百步内遗忘状态脉冲,因为每次写入都会沿当前键方向覆盖状态;(iv)每token量化成本随上下文抵消而非累积。我们还修复了一个全局规模不匹配问题:当按模块校准的NVFP4检查点被内核融合为单个GEMM时会出现该问题,并证明校准后的FP8键值(KV)缓存缩放无性能损失。最终得到实用方案:量化所有内容,随模型发布KV缩放,同时阐明了混合LLM的循环部分为何是量化难度较低的部分。检查点地址:this https URL

英文摘要

Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent state summarizes the context in fixed size. Early community 4-bit quantizations of Qwen3.8-27B (48 GDN layers, 16 attention layers) left the GDN block in 8- or 16-bit precision -- especially its decay and write-strength gates -- on the intuition that errors in a recurrence accumulate over long contexts. We test that intuition by building Minima: NVFP4 W4A4 on all 496 linear layers, GDN included. Across perplexity at 4K/32K, MMLU-Pro, GSM8K, AIME'25, GPQA-Diamond, LiveCodeBench, and RULER retrieval to 64K, Minima matches BF16 within seed noise (5-task average -0.52) while being the smallest (17.5 GiB) and fastest-prefill (+14-19%) recipe we compare, and its 32K perplexity gap shrinks with position. A four-part mechanism study explains why: (i) NVFP4's 16-element block scaling localizes the residual stream's extreme outliers, equalizing activation error across layer roles; (ii) the supposedly fragile gate projections are the least sensitive -- softplus/exponential and sigmoid parameterizations compress ~11% GEMM error to ~2% output error; (iii) the delta-rule recurrence holds injected noise at a flat plateau over 32K tokens and forgets a state impulse within hundreds of steps, because each write overwrites the state along the current key direction; (iv) the per-token quantization cost washes out with context instead of compounding. We also repair a global-scale mismatch that arises when per-module-calibrated NVFP4 checkpoints are served by kernels that fuse those modules into one GEMM, and show calibrated FP8 KV-cache scales are performance-free. The result: a practical recipe -- quantize everything, ship KV scales -- and a mechanistic account of why the recurrent half of a hybrid LLM is the easy half to quantize. Checkpoint: https://huggingface.co/minima-ai/mnma_qwen3.8_27b_nvfp4

Comments14 pages, 2 figures, 6 tables. Quantized checkpoint: https://huggingface.co/minima-ai/mnma_qwen3.8_27b_nvfp4

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑