arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24799cs.CLcs.AI

当量化保留准确性但不保留证据时:面向医学大语言模型的解释感知训练后量化

When Quantization Preserves Accuracy but Not Evidence: Explanation-Aware Post-Training Quantization for Medical LLMs

Yeji Kim, Mi-Young Kim, Randy Goebel

首次发表
浏览论文内容

中文总结 AI 辅助

针对医学LLM的PTQ,提出解释感知目标,通过保真度缓存保留支持答案的证据,优于仅保留准确性的基线。

中文摘要 AI 辅助

训练后量化(PTQ)能够实现大语言模型的高效部署,且PTQ方法通常以通用重建、困惑度或答案准确性进行优化和评估。但在解释关键型领域,仅保留最终答案可能不足,因为用户可能还会检查生成的推理过程以判断预测是否可信。我们在医学多项选择题回答中研究此问题,其中推理过程应提供支持所选答案的证据。我们提出了一种面向基于变换的PTQ的解释感知目标。我们的方法从全精度教师推理过程中构建离线保真度缓存,并在优化期间使用它来保留支持答案的证据标记和证据条件下的答案行为。我们在W4A4KV4量化下于OSTQuant上实例化该方法,并在MedExQA、MedExpQA和ChallengeClinicalQA上评估了四个7B-8B医学和指令调优的大语言模型。虽然相同校准的OSTQuant基线保留了任务准确性,但它可能显著削弱支持答案的推理过程。我们的目标是保留全精度模型的答案支持行为,而非提高黄金标签准确性,且我们的方法更好地保留了全精度模型的答案行为和推理到答案的支持。这些结果表明,解释关键型设置的PTQ应评估答案支持证据的保留,而不仅仅是答案准确性。代码和评估脚本可在以下网址获取:此https URL。

英文摘要

Post-training quantization (PTQ) enables efficient deployment of large language models, and PTQ methods are usually optimized and evaluated with generic reconstruction, perplexity, or answer accuracy. But in explanation-critical domains, preserving only the final answer may be insufficient, since users may also inspect generated rationales to judge whether a prediction is trustworthy. We study this issue in medical multiple-choice question answering, where rationales should provide evidence that supports the selected answer. We propose an explanation-aware objective for transformation-based PTQ. Our method builds an offline faithfulness cache from full-precision teacher rationales and uses it during optimization to preserve answer-supporting evidence tokens and evidence-conditioned answer behavior. We instantiate it on OSTQuant under W4A4KV4 quantization and evaluate four 7B--8B medical and instruction-tuned LLMs on MedExQA, MedExpQA, and ChallengeClinicalQA. While a same-calibration OSTQuant baseline preserves task accuracy, it can substantially weaken answer-supporting rationales. Our objective is to preserve the full-precision model's answer-supporting behavior rather than improve gold-label accuracy, and our method better preserves the full-precision model's answer behavior and rationale-to-answer support. These results suggest that PTQ for explanation-critical settings should evaluate preservation of answer-supporting evidence, not only answer accuracy. Code and evaluation scripts are available at https://github.com/dut0817/EAQuant.

发表机构

  • University of Alberta(阿尔伯塔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑