面向量化语言模型不确定性保留的目标感知校准数据选择
Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models
浏览论文内容
中文总结 AI 辅助
本文针对量化语言模型的不确定性保留问题,提出DPQ方案,通过选择与目标对齐的校准数据,在多模型多基准实验中验证其能根据部署需求保留不同类型的不确定性行为。
中文摘要 AI 辅助
量化技术被广泛用于部署大型语言模型,但其对置信度、决策边界、弃权(不执行)等不确定性行为的影响很少被作为核心目标。本文将面向量化的校准数据选择问题建模为依赖目标的不确定性保留问题:不同部署场景侧重输入分布的不同区域,而现有工作主要优化面向准确率的压缩指标,或在量化后调整分数。本文通过分布保留风险与边界保留风险形式化该目标,并提出简单的分布不匹配混合论,解释不存在适用于所有目标的单一校准方案。本文引入轻量级预量化方案家族Doubt-Preserving Quantization(DPQ),利用全精度预测结果构建与目标对齐的校准混合集,包含高置信度示例与通用锚点。在8个语言模型、9个NLP基准、22种对比方法的实验中,最优固定方案随保留目标变化:DPQ-r75在SQuAD2的答案边界保留上表现最优,而更温和或单信号变体(含DPQ-r50、仅置信度、仅熵)更适合保留广泛的多选QA行为。这些结果表明,校准数据应根据部署所需保留的特定全精度分数行为进行选择,而非作为固定的量化细节处理。
英文摘要
Quantization is widely used to deploy large language models, but its effect on uncertainty behavior, such as confidence, margins, and abstention, is rarely treated as a primary objective. We frame calibration-data selection for quantization as a target-dependent uncertainty-preservation problem. Different deployments emphasize different regions of the input distribution, yet prior work mainly optimizes accuracy-oriented compression metrics or adjusts scores after quantization. We formalize this goal with distributional and boundary preservation risks, and provide a simple mixture-mismatch argument explaining why no single calibration recipe should be expected to fit all targets. We introduce Doubt-Preserving Quantization (DPQ), a lightweight pre-quantization recipe family that uses full-precision predictions to construct target-aligned calibration mixtures of high-doubt examples and generic anchors. Across 8 language models, 9 NLP benchmarks, and 22 comparison methods, the leading fixed recipe changes with the preservation target: DPQ-r75 leads on SQuAD2 answerability-boundary preservation, while milder or single-signal variants, including DPQ-r50, confidence-only, and entropy-only, better preserve broad multiple-choice QA behavior. These results show that calibration data should be selected for the specific full-precision score behavior a deployment needs to preserve, rather than treated as a fixed quantization detail.
发表机构
- The Hong Kong University of Science and Technology(香港科技大学)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- City University of Hong Kong(香港城市大学)
- Shanghai Institute of Optics and Fine Mechanics(上海光学精密机械研究所)
- University of Oxford(牛津大学)
- Yale University(耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。