arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量化对临床基准的影响:不同模型家族的准确性与安全性

The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families

Leonard Twagirayezu, Prasenjit Mitra

arXiv 2609.22216首次发表:更新:

发表机构

Carnegie Mellon University Africa(卡内基梅隆大学非洲校区)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估了不同模型家族在FP16、INT8和INT4量化下的临床基准表现,发现INT8普遍安全,INT4退化显著且依赖模型和任务,安全对齐主要由指令微调决定。

AI 中文摘要

量化使得大型语言模型能够在资源受限的临床边缘设备上部署,但其对临床准确性和安全性的影响仍未得到充分研究。我们在五个基准上评估了五个7-8B参数模型在FP16、GPTQ-INT8和GPTQ-INT4精度下的表现:MedQA、MedMCQA、Med-HALT、HealthBench的风险分层样本以及MedSafetyBench。该研究联合变化量化位宽、模型家族和临床任务类型,并明确进行风险分层和安全措施。INT8 GPTQ普遍安全(最大退化-1.9%至1.9%),而INT4退化显著且依赖模型:经过临床微调的BioMistral-7B在MedMCQA上损失19.7%,超过任何通用模型,表明临床微调并未赋予压缩鲁棒性。在INT4下,MedMCQA的退化程度高于MedQA;Med-HALT基本不受影响。在HealthBench的紧急风险子组中,Qwen2.5-7B在INT4下退化26.8%,表明高风险场景对压缩尤为脆弱。在MedSafetyBench上,模型家族的影响超过精度(FP16下拒绝率范围为10.2%-74.9%),尽管Qwen2.5-7B(-17.8%)和Meditron-7B(-28.3%)表现出显著的INT4安全退化;值得注意的是,Qwen2.5-7B同时是最具准确性鲁棒性的模型,表明准确性和安全性鲁棒性是独立的属性。我们还测试了两种恢复方法:临床校准替换和QLoRA微调,两者产生相同的权衡:MedMCQA恢复而MedQA进一步退化,表明恢复策略需要针对特定任务进行验证,而不能假定其普遍有益。这些发现表明INT8对于临床部署总体安全,而INT4的安全性必须按模型和任务进行评估,且安全对齐主要由指令微调而非临床领域适应决定。

英文摘要

Quantization enables deployment of large language models on resource-constrained clinical edge devices, but its effect on clinical accuracy and safety remains understudied. We evaluate five 7-8B parameter models at FP16, GPTQ-INT8, and GPTQ-INT4 precision across five benchmarks: MedQA, MedMCQA, Med-HALT, a risk-stratified sample of HealthBench, and MedSafetyBench. The study jointly varies quantization bit width, model family, and clinical task type, with explicit risk stratification and safety measures. INT8 GPTQ is universally safe (max. degradation -1.9%-1.9%), while INT4 degradation is substantial and model-dependent: BioMistral-7B, clinically fine-tuned, loses 19.7% on MedMCQA, more than any general-purpose model, showing clinical fine-tuning does not confer compression robustness. MedMCQA degrades more than MedQA under INT4; Med-HALT is largely unaffected. On HealthBench's emergency-risk subgroup, Qwen2.5-7B degrades by 26.8% under INT4, suggesting high-risk scenarios are disproportionately vulnerable to compression. On MedSafetyBench, the model family dominates over precision (refusal rates range 10.2%-74.9% at FP16), though Qwen2.5-7B (-17.8%) and Meditron-7B (-28.3%) show substantial INT4 safety degradation; notably, Qwen2.5-7B is simultaneously the most accuracy-robust model, demonstrating that accuracy and safety robustness are independent properties. We additionally test two recovery methods, clinical calibration substitution and QLoRA fine-tuning, both producing the same trade-off: MedMCQA recovers while MedQA further degrades, indicating recovery strategies require task-specific validation rather than being assumed universally beneficial. These findings indicate INT8 is broadly safe for clinical deployment, while INT4 safety must be assessed per-model and per-task, and that safety alignment is determined primarily by instruction tuning rather than clinical domain adaptation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑