arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23585cs.LGcs.CL

全局排名保持,选定的注意力头发生偏移:4比特仅权重量化下的BOS-Sink拓扑

Global Ranks Survive, Selected Heads Shift: BOS-Sink Topology under 4-bit Weight-Only Quantization

Kuanlin Chen, Chen-Wei Kuo, Cheng-En Ou

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过Sink拓扑一致性指标评估4比特仅权重量化对首token注意力头的影响,发现全局排名保持但离散头集和跨域校准需重新验证。

中文摘要 AI 辅助

Sink感知部署可以在模型量化之前识别重要的首token注意力头,然后在边缘设备上重用该映射。我们测试了这种捷径在4比特NF4仅权重训练后量化(PTQ)中何时是安全的。我们的Sink拓扑一致性(STC)指标分别衡量全局排名保持、top-$k$集合重叠和逐层sink质量偏移,并区分逐输入敏感性与校准映射迁移。在Qwen2.5-0.5B、Qwen2.5-1.5B和Llama-3.2-1B上,全局bf16到4比特排名在4,096个token时保持较高($\rho_s \geq 0.980$),但top-$k$ Jaccard重叠仅为0.619-0.793,对应76.5-88.5%的成员保留率。全局统计量也掩盖了局部失败:Qwen的终端层偏移量是其模型均值的6.2-7.9倍,而Llama-3.2-1B显示出低且几乎均匀的漂移。在C4到LongBench的域转移下,两个Qwen模型的跨域重叠比域内精度比较下降更多,但Llama-3.2-1B并非如此。匹配域的4比特重新校准在最小测试$n=8$时达到分割半稳定性平台的90%,对于两个Qwen模型,而Llama-3.2-1B在$n=32$时达到,但并非作为尖锐阈值;对于两个Qwen模型,仅更新选定的层无法达到全映射稳定性标准。在Jetson Orin NX上,16样本工作负载对于具有有效设备上sink测量的两个模型只需数秒。实际信息是明确的:全局排名通常可迁移,但离散的注意力头集合、层局部策略和跨域校准应在量化后重新验证。

英文摘要

Sink-aware deployment may identify important first-token attention heads before a model is quantized, then reuse that map at the edge. We test when this shortcut is safe for 4-bit NF4 weight-only post-training quantization (PTQ). Our Sink Topology Consistency (STC) metrics separate global rank preservation, top-$k$ set overlap, and layerwise sink-mass shift, and distinguish per-input sensitivity from calibration-map transfer. Across Qwen2.5-0.5B, Qwen2.5-1.5B, and Llama-3.2-1B, global bf16-to-4-bit ranks remain high at 4,096 tokens ($ρ_s \geq 0.980$), yet top-$k$ Jaccard overlap is only 0.619-0.793, corresponding to 76.5-88.5% membership retention. The global statistic also masks local failures: terminal Qwen layers shift by 6.2-7.9x their model means, whereas Llama-3.2-1B shows low, nearly uniform drift. Under a C4-to-LongBench shift, cross-domain overlap degrades more than the within-domain precision comparison for both Qwen models, but not for Llama-3.2-1B. Matched-domain 4-bit recalibration reaches 90% of a split-half stability plateau at the smallest tested $n=8$ for both Qwen models and $n=32$ for Llama-3.2-1B, though not as a sharp threshold; for the two Qwen models, updating only selected layers does not reach the full-map stability criterion. On Jetson Orin NX, the 16-sample workload takes seconds for the two models with valid on-device sink measurements. The practical message is precise: global rankings often transfer, but discrete head sets, layer-local policies, and cross-domain calibration should be revalidated after quantization.

发表机构

  • National Tsing Hua University(国立清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑