发表机构
University of Electronic Science and Technology of China(电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究医学大语言模型诊断时证据使用情况,通过分解患者信息为证据单元等方法审计,在多个数据集上评估五个模型,发现准确性可能掩盖证据使用问题,结果促使对医学大语言模型评估进行角色感知审计。
AI 中文摘要
医学大语言模型通常根据是否选择正确诊断来评估,但仅诊断准确性并不能表明模型是否正确使用了病例证据。我们对医学诊断中的证据使用进行了行为审计。对于每个病例,我们将患者信息分解为证据单元,在受控证据子集中对候选诊断进行评分,并挖掘诊断边际中的低阶相互作用。由于医学证据与诊断相关,审计将相互作用发现与失败分配分开:大的或负的相互作用可反映合理的鉴别诊断,而可疑的相互作用需要进行稳健性检查和临床审查。我们在DDXPlus、CupCase和MedCase上评估了五个开放权重的大语言模型。在各个数据集中,忠实支持和鉴别冲突或抵消占大多数相互作用强度,表明许多证据相互作用在临床上是合理的而非失败。在一个以DDXPlus为重点的由五名审阅者进行的130项盲法丰富审阅样本中,无效或类似捷径的病例集中在否定或缺失的发现以及临床局部证据中。这些结果表明,准确性可能掩盖候选证据使用失败的情况,并促使对医学大语言模型评估进行角色感知审计。
英文摘要
Medical LLMs are often evaluated by whether they select the correct diagnosis, but diagnostic accuracy alone does not show whether the model used the case evidence appropriately. We present a behavioral audit of evidence use in medical diagnosis. For each case, we decompose patient information into evidence units, score candidate diagnoses under controlled evidence subsets, and mine low-order interactions in diagnostic margins. Because medical evidence is diagnosis-relative, the audit separates interaction discovery from failure assignment: large or negative interactions can reflect plausible differential diagnosis, while suspicious interactions require robustness checks and clinical review. We evaluate five open-weight LLMs on DDXPlus, CupCase, and MedCase. Across datasets, faithful support and differential conflict or cancellation account for most interaction strength, showing that many evidence interactions are clinically plausible rather than failures. In a DDXPlus-focused blinded five-reviewer 130-item enriched review sample, invalid or shortcut-like cases concentrate in negated or absent findings and clinically local evidence. These results show that accuracy can hide candidate evidence-use failures and motivate role-aware audits for medical LLM evaluation.