开源权重大语言模型中的临床风险分类错位与成本不对称性
Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models
- Johns Hopkins University School of Medicine(约翰斯·霍普金斯大学医学院)
- Microsoft(微软)
- National Institutes of Health(美国国立卫生研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究揭示开源大语言模型虽能编码患者风险与成本信息,但未能可靠整合以做出成本正确的临床决策,且存在方向与幅度处理的不对称性,提示临床评估需补充权衡测试。
AI中文摘要:
大语言模型(LLMs)如何将患者风险与临床成本权衡相结合,目前仍知之甚少。我们研究了四个开源权重LLMs(Qwen-2.5-7B/32B和Llama-3.1-8B/70B)如何在内部表示成本权衡,这些表示如何与临床预测相关联,以及决策是否按照指定的成本方向和幅度所预测的方式发生改变。利用一个公共糖尿病数据集,我们在三种措辞下变化了11种假阴性(FN)与假阳性(FP)成本比率,并检查了表示和行为的输出。患者风险可线性恢复,其性能与常规分类器相当(AUC约0.83),且成本方向在每个模型中均可恢复。然而,成本方向的表示变化仅在两个较大的模型中跟踪了输出变化,而对成本幅度的响应主要与方向无关。在12个模型-措辞组合中,仅有2个表现出对FN与FP成本增加的对立响应以及成本正确的排序。在表示层面,在一个成本侧拟合的方向转移到另一侧时并未反转,这与镜像对称编码下的预期不符。这些发现表明,LLMs编码了风险和成本信息,但并未可靠地将它们整合到成本正确的决策中。因此,临床评估应在预测性能之外,纳入权衡测试、措辞敏感性和默认操作点。
英文摘要:
How large language models (LLMs) integrate patient risk with clinical cost tradeoffs remains poorly understood. We investigated how four open-weight LLMs (Qwen-2.5-7B/32B and Llama-3.1-8B/70B) internally represent cost tradeoffs, how these representations relate to clinical predictions, and whether decisions shift as predicted by the specified cost direction and magnitude. Using a public diabetes dataset, we varied 11 false-negative (FN) to false-positive (FP) cost ratios across three phrasings and examined representations and behavioral outputs. Patient risk was linearly recoverable on par with conventional classifiers (AUC $\approx 0.83$), and cost direction was recoverable in every model. However, representational shifts in cost direction tracked output changes only in the two larger models, and responses to cost magnitude were predominantly direction-agnostic. Only 2 of 12 model-phrasings showed both opposing responses to increasing FN versus FP costs and cost-correct ordering. Representationally, a direction fitted on one cost side did not invert when transferred to the other, as expected under mirror-symmetric encoding. These findings suggest that LLMs encode risk and cost information but do not reliably integrate them into cost-correct decisions. Clinical evaluations should therefore include tradeoff tests, phrasing sensitivity, and default operating points alongside predictive performance.