评估药物安全推理中对患者信息的反事实敏感性
Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning
AI总结:
本研究推出MedPIC-Bench基准,评估不同LLM在药物安全推理中对患者信息的反事实敏感性,发现所有模型在反事实问题上表现更差,医学专用LLM的反事实表现落后于通用LLM,凸显静态准确性的评估局限性。
AI中文摘要:
当药物安全规则的患者特定条件未满足时应用该规则会产生错误决策,现有医学评估大多采用孤立且固定的场景,模型可能通过回忆药物-风险关联得出正确答案,却未体现其利用患者信息判断规则是否适用。为解决该缺口,我们推出MedPIC-Bench,这是一个包含来源可验证建议和专家验证问题的基准,用于患者特定药物安全推理,它将遵循指南的问题与配对的反事实问题相结合,其中患者信息的受控变化会改变规则是否适用。该基准包含467个问题,沿6个临床和推理维度标注。在28个医学专用、通用及专有大型语言模型(LLM)中,所有模型在反事实问题上的表现更差,平均准确率从63.6%降至45.1%。当明确的患者属性直接表明熟悉的禁忌症时,模型表现良好;但当患者信息需缩小或撤销安全警告时,模型会陷入困境。模型的理由常承认患者信息的变化,但最终答案仍保留之前的安全判断,这种脆弱性在医学专用LLM中依然存在,其平均反事实(CF)表现落后于通用LLM。因此,MedPIC-Bench使条件规则应用可测量,并凸显了静态药物安全准确性在评估患者特定可靠性方面的局限性。
英文摘要:
Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations largely use isolated and fixed scenarios. A model may therefore answer correctly by recalling a drug-risk association without showing that it used patient information to decide whether the rule applies. To address this gap, we introduce MedPIC-Bench, a benchmark of source-verifiable recommendations and expert-validated questions for patient-specific medication-safety reasoning. It combines guideline-following questions with paired counterfactual questions in which a controlled change in patient information changes whether a rule applies. The benchmark contains 467 questions annotated along six clinical and reasoning dimensions. Across 28 medical-specific, general, and proprietary LLMs, every model performs worse on counterfactual questions, with mean accuracy falling from 63.6\% to 45.1\%. Models perform well when an explicit patient attribute directly signals a familiar contraindication, but struggle when patient information must narrow or withdraw a safety warning. Model rationales often acknowledge the changed patient information, yet the final answers retain the previous safety judgment. This vulnerability persists among medical-specific LLMs, whose average CF performance trails that of general LLMs. MedPIC-Bench therefore makes conditional rule application measurable and highlights the limitations of static medication-safety accuracy for assessing patient-specific reliability.