arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型中的临床概念中心

Clinical Concept Centers in LLMs

Aishik Nagar, Abhishek Vaidyanathan, Arun-Kumar Kaliya-Perumal, Elijah Tzen Hsuen Boey, Stefan Winkler

arXiv 2610.02829首次发表:更新:

发表机构

National University of Singapore; SAP Asia Pte Ltd; Nanyang Technological University; Singapore Institute of Technology(新加坡国立大学; SAP亚洲私人有限公司; 南洋理工大学; 新加坡理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将行为评估扩展到大语言模型的潜在空间,发现开放权重模型中存在可解释且因果驱动行为的临床概念中心,并验证其在评估、性能提升及临床偏好预测方面的价值。

AI 中文摘要

大语言模型在临床环境中的应用日益广泛。然而,关于这些模型的可靠性和性能的研究几乎完全集中在语言层面,即对模型所说内容进行评分。机制可解释性研究发现,潜在空间比文本具有更高的表征保真度:内部表征不仅编码了远超输出所表达的内容,而且所述推理也系统性地遗漏了因果驱动答案的特征。在临床决策支持中,尚未探索基于机制可解释性的模型行为评估。在本工作中,我们将行为评估扩展到潜在空间,并探究临床概念是否以可定位、因果使用的方式存在于开放权重大语言模型内部的表征中。我们在所测试的全部十一个开放模型的潜在空间中都发现了专门的临床概念中心。这些概念中心是可解释的,仅在其对应的临床叙述中激活,并在受约束和开放式环境中均有意义地且因果性地驱动模型行为。它们不仅仅是分析性表征,更是可在临床实践中利用的电路,我们从评估和性能两个角度探讨了其用途。从评估角度看,即使在对抗性角色基元提示下,模型仍保持内部一致性并持续使用相关概念中心,而对齐基元提示则改善了下游临床性能。从性能角度看,我们模拟了现实部署场景,发现沿这些中心对模型进行引导可带来有意义的下游改进。最后,我们进行了盲法临床医生验证,发现这些概念中心的激活和使用能够预测临床医生的偏好。

英文摘要

Large language models are increasingly used in clinical settings. However, research into the reliability and performance of these models has focused almost entirely on the language substrate, scoring what the model says. Mechanistic interpretability has found that the latent space carries a higher fidelity of representation than the text: internal representations not only encode substantially more than the output verbalizes, but the stated reasoning also systematically omits features that causally drive the answer. An evaluation of model behavior in terms of mechanistic interpretability has not been explored in clinical decision support. In this work, we extend behavioral evaluation into the latent space and ask whether clinical concepts exist as locatable, causally used representations inside open-weight LLMs. We find dedicated clinical concept centers in the latent space of all eleven open models we test. These concept centers are interpretable, firing only on their aligned clinical narratives, and meaningfully and causally drive model behavior in both constrained and open-ended settings. They are not just analytical representations, but circuits that can be utilized in clinical practice, and we explore their use from the perspective of both evaluation and performance. From the evaluation standpoint, models stay internally coherent and keep using the relevant concept centers even under adversarial role-based priming, while aligned priming improves downstream clinical performance. From a performance perspective, we simulate realistic deployment settings and find that steering models along these centers leads to meaningful downstream improvements. Finally, we conduct a blinded clinician validation and find the activation and usage of these concept centers predicts clinicians preferences.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑