arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27397cs.CLcs.AI

使临床语言模型可审计:概念引导的鲁棒预测微调

Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction

Jin Mu, Guanhua Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出基于SAE的可审计临床文本分类框架CAST,通过提取并抑制病历人工制品特征,在MIMIC-IV死亡率预测任务上提升了基线性能并保留与强LLM相当的效果,同时提供可审计的特征级轨迹。

中文摘要 AI 辅助

临床语言模型在院内预测中可达到较高准确率,但在部署环境发生变化时会失效,因为它们利用了病历特有的人工制品(如模板、分隔符、套话),这些人工制品并不反映患者状态。我们提出CAST(Concept-guided Artifact Suppression Tuning,概念引导的人工制品抑制微调),一种基于SAE的可审计临床文本分类框架。CAST使用稀疏自编码器(Sparse Autoencoders)从Transformer中间激活中提取稀疏、可人工审计的特征,通过大语言模型辅助的解释流程和ICD-10检索约束对SAE隐变量进行标注,在微调阶段通过残差减法抑制已验证的人工制品隐变量,并提供事后逐概念归因以审计模型决策。在MIMIC-IV出院 note 死亡率预测任务上,CAST优于对应的微调编码器基线,且与强大的大语言模型基线性能相当,同时生成支持每个预测的临床概念以及训练期间被抑制的人工制品概念的特征级审计轨迹。

英文摘要

Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifacts (e.g., templates, separators, boilerplate) that do not reflect patient state. We propose CAST (Concept-guided Artifact Suppression Tuning), an SAE-based framework for auditable clinical text classification. CAST uses Sparse Autoencoders to expose sparse, human-auditable features from intermediate Transformer activations, labels SAE latents with an LLM-assisted interpretation pipeline and ICD-10 retrieval constraints, suppresses verified artifact latents via residual subtraction during fine-tuning, and provides post-hoc per-concept attributions for auditing model decisions. On MIMIC-IV discharge-note mortality prediction, CAST improves over its corresponding fine-tuned encoder baselines and remains competitive with strong LLM baselines, while producing a feature-level audit trail of the clinical concepts that support each prediction and the artifact concepts suppressed during training.

发表机构

  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

↑