arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量化对生物医学大语言模型可靠性的影响

When Calibration Depends on the Scoring Rule: Quantized Biomedical LLM Classification

Anton Rasmussen, Hong Qin

arXiv 2608.03854首次发表:更新:

AI 中文总结

该研究针对 Mistral-7B 变体在 PubMed RCT 句子分类任务上,评估量化及提示模板、评分规则等对生物医学大语言模型可靠性的影响,发现这些因素对校准和准确率的影响不可忽视。

AI 中文摘要

当解码器语言模型用作分类器时,预测的类别概率取决于实现选择,包括提示模板、 verbalizer(标签到标记的映射)和评分规则,这些很少被视为实验变量。我们在 FP16、INT8 和 INT4 精度下,使用四种答案文本提示模板,对三个 Mistral-7B 变体(Base、BioMistral 和 Instruct)在 PubMed RCT 句子分类(n=2000)上进行了受控评估。我们的主要发现是,概率提取协议主导了表观校准。从求和标记对数似然评分切换为均值标记对数似然评分会反转模型间的校准排名:BioMistral 的平均预期校准误差从 0.097 增加到 0.289,而 Instruct 的平均预期校准误差从 0.237 降低到 0.096;对于专用模型,准确率变化小于 1 个百分点,而基础模型的准确率变化为 4-6 个百分点。提示模板选择产生的准确率差异为 7-24 个百分点,与模型级效应相当或更大。在某一模板上,BioMistral 的表现优于 Instruct,尽管总体均值仅使 Instruct 领先 1.3 个百分点。对于 BioMistral 和 Instruct,INT8 量化相对于 FP16 仅使准确率和 F1 变化 1-2 个百分点,而基础模型在某些模板上显示出更大的 INT8 效应(高达 +4.2 个百分点)。INT4 产生了异质性但非灾难性的效应。温度缩放可降低两种模型在求和评分下的预期校准误差,但仅针对该评分规则。一个微调的 PubMedBERT 参考模型达到 82.7% 的准确率,但使用了约 176000 个带标签的训练示例,无法直接比较。这些结果表明,在评估解码器语言模型校准时,提示模板设计和评分归一化是一级实验决策。

英文摘要

Quantized large language models can run on consumer hardware, which motivates interest in on-premises processing of sensitive data. The reliability of their confidence estimates depends on implementation choices (prompt template, label wording, and scoring normalization) that are seldom treated as experimental variables. We evaluate three 7-billion-parameter Mistral variants (a base model, BioMistral, and an instruction-tuned checkpoint evaluated without its chat template) at FP16, INT8, and INT4 on five-class sentence classification in medical abstracts. We test two primary templates on n = 2,000 test sentences and two auxiliary templates on n = 200 validation sentences. Our central observation, made without post-hoc calibration, is that switching from summed to mean-token log-likelihood reverses which model appears better calibrated: BioMistral's mean calibration error across matched conditions nearly triples, while the instruction-tuned model's drops by more than half. Accuracy changes by at most 1.4 percentage points for these two checkpoints. Negative log-likelihood and Brier score show the same reversal. The two primary templates were selected using test-derived examples, so absolute performance with them is exploratory. Between them, prompt choice changes mean accuracy across precisions by 2.9 to 17.8 percentage points. Eight-bit quantization changes accuracy by at most 1.1 percentage points for the adapted checkpoints; four-bit quantization shows mixed but non-catastrophic effects. Post-hoc temperature scaling reduces calibration error under summed scoring but was not fitted under mean-token scoring, so whether the reversal survives per-scorer calibration is unknown. These exploratory results suggest that calibration comparisons of decoder-based classifiers should treat scoring normalization and prompt design as first-order experimental decisions.

Comments8 pages, 1 figure. Accepted at the AAAI 2026 Fall Symposium Series (AT-AI4H-NW 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑