arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09999cs.CLcs.LG

量化语言模型推理中的无声失败:基于分类法的空洞收敛和失败模式转移分析

Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts

Renuka Oladri, Mohan Vamsi Varadaraju Priya, Jerry Wu

首次发表
浏览论文内容

中文总结 AI 辅助

研究量化大语言模型推理中的无声失败,用六类失败分类法对五个模型在三种精度和四个基准下的思维链输出分类,发现空洞收敛等问题有精度和基准依赖性,且无法从文本特征可靠检测,揭示了标准评估管道未捕捉到的失败模式。

中文摘要 AI 辅助

我们表明,即使任务准确性得以保持,训练后量化也会悄然改变大语言模型的推理方式。我们使用由两名独立人工标注者验证的六类失败分类法(科恩卡方系数κ = 0.906),对五个指令微调的大语言模型(参数为3B - 14B)在三种量化精度(FP32、FP16、NF4)和四个推理基准下的30,000个思维链输出进行分类。我们发现,虽然跨精度的准确性较为稳健(最大下降3.1个百分点),但空洞收敛(通过不完整或不可验证的推理得出正确答案)在NF4下显示出显著的规模依赖性转移,对于测试的两个最小模型急剧下降,而对于12B参数及以上的模型保持不变。这种影响也是特定于基准的:GSM8K完全免疫,而LogiQA和ARC - Challenge显示出最大的转移。此外,在NF4下,捷径崩溃在LLaMA 3.2 - 3B中从错误答案失败的44%上升到78%,而信心雪球效应从15.8%降至接近零,这是准确性指标无法察觉的定性转移。最后,我们表明无法从表面文本特征可靠地检测到空洞收敛(最佳F1 = 0.53),将其确立为标准评估管道无法捕捉的与部署相关的失败模式。

英文摘要

We show that post-training quantization can silently alter how large language models reason even when task accuracy is preserved. Using a six-category failure taxonomy validated by two independent human annotators (Cohen's $κ$ = 0.906), we classify 30,000 chain-of-thought outputs from five instruction-tuned LLMs (3B--14B parameters) across three quantization precisions (FP32, FP16, NF4) and four reasoning benchmarks. We find that while accuracy is robust across precisions (maximum 3.1 pp drop), Hollow Convergence (correct answers reached through incomplete or unverifiable reasoning) shows a significant size-dependent shift under NF4, dropping sharply for the two smallest models tested but remaining invariant for models at 12B parameters and above. This effect is also benchmark-specific: GSM8K is categorically immune while LogiQA and ARC-Challenge show the largest shifts. Furthermore, under NF4, Shortcut Collapse rises from 44% to 78% of wrong-answer failures in LLaMA 3.2-3B while Confidence Snowballing collapses from 15.8% to near zero, a qualitative shift invisible to accuracy metrics. Finally, we show Hollow Convergence cannot be reliably detected from surface-level text features (best F1 = 0.53), establishing it as a deployment-relevant failure mode that standard evaluation pipelines cannot catch.

↑