发表机构
Institute of Product Leadership; Amazon, Inc(产品领导力学院; 亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对临床LLM解释忠实度问题,提出ESI、CFS和PSS三个轻量指标,实验发现仅23.3%引用概念因果必要,且准确率、一致性与共识均不能可靠反映推理忠实度。
AI 中文摘要
临床大型语言模型(LLMs)在医学考试中取得了较高的准确性;然而,正确的答案并不能保证解释中提及的概念确实是驱动决策的因素。我们引入了三个轻量级、可直接解释的指标来衡量这种忠实度差距:解释稳定性指数(ESI),用于衡量重复查询中推理的一致性;因果忠实度评分(CFS),通过概念消融测试被引用的概念是否驱动预测;以及扰动稳定性评分(PSS),用于衡量对语义保持改写(paraphrases)的鲁棒性。通过在150道MedQA-USMLE问题(900个模型-问题观测)上评估六个LLM,我们发现仅有23.3%的被引用临床概念是因果上必要的。正确回答的CFS低于错误回答(0.212对0.398),回答一致性对CFS呈负预测(Spearman r = -0.466),并且模型对可能在答案上一致,但只共享8.8%的被引用推理概念。这些结果表明,准确性、一致性和共识对于临床决策支持而言是不完整的安全信号。证据是行为性的而非机制性的:概念消融测试的是输出的反事实敏感性,而非内部电路。
英文摘要
Clinical Large Language Models (LLMs) achieve strong medical-exam accuracy; however, a correct answer does not guarantee that the explanation names the concepts that actually drove the decision. We introduce three lightweight, directly interpretable metrics for this faithfulness gap: the Explanation Stability Index (ESI), which measures reasoning consistency across repeated queries; the Causal Faithfulness Score (CFS), which tests whether cited concepts drive predictions via concept ablation; and the Perturbation Stability Score (PSS), which measures robustness to semantic-preserving paraphrases. By evaluating six LLMs on 150 MedQA-USMLE questions (900 model-question observations), we found that only 23.3% of the cited clinical concepts were causally necessary. Correct answers had lower CFS than incorrect answers (0.212 vs. 0.398), answer consistency negatively predicted CFS (Spearman r = -0.466), and model pairs could agree on answers while sharing only 8.8% of cited reasoning concepts. These results show that accuracy, consistency, and consensus are incomplete safety signals for clinical decision-making support. The evidence is behavioral rather than mechanistic: concept ablation tests counterfactual sensitivity of outputs, not internal circuits.
CommentsAccepted and received best paper award from ICETCI 2026 Conference, https://ietcint.com/