诊断前先询问:Safe-Psych,一种用于精神病学领域大语言模型的顺序评估基准
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
浏览论文内容
中文总结 AI 辅助
研究针对大语言模型在精神病学决策支持中处理证据不完整问题,引入顺序评估基准Safe-Psych,通过模拟真实临床记录及精神科医生行动标签,评估多个模型,发现模型存在过早诊断等问题,揭示其局限性,为改善模型安全性研究提供支持。
中文摘要 AI 辅助
大语言模型(LLMs)越来越多地用于医疗保健决策支持,但临床证据往往不完整或不断变化。当可用信息不足以支持可靠答案时,模型应要求澄清或弃权,而不是提供无根据的回答。然而,现有的医学基准通常假设一开始就有完整信息。我们引入了Safe-Psych,这是一个用于评估LLMs如何处理临床精神病学中不断变化的诊断不确定性的顺序基准。Safe-Psych包含1000多个真实世界的精神病临床记录,这些记录被分割以模拟渐进式证据披露,每个阶段都有精神科医生给出的行动标签:诊断、澄清或弃权。我们在全信息和顺序设置中评估了多个先进的LLMs。研究结果表明,能力并不能确保校准:即使是强大的模型在临床信息不完整的情况下也会挣扎,大多数模型的弃权不足超过60%,安全意识提示只是将错误转向过度弃权,从而减少过早承诺。在顺序评估中,模型经常在没有足够证据的情况下进行诊断,除非明确提示,否则很少寻求澄清;这些过早诊断的准确性低于及时诊断。总体而言,Safe-Psych揭示了所评估模型的一个局限性:识别临床证据何时不完整以及何时需要额外信息。我们发布Safe-Psych以支持改善医疗保健中LLM安全性的研究。
英文摘要
Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving. When the available information is insufficient to support a reliable answer, models should request clarification or abstain rather than provide unsupported responses. Existing medical benchmarks, however, typically assume that complete information is available upfront. We introduce Safe-Psych, a sequential benchmark for evaluating how LLMs handle evolving diagnostic uncertainty in clinical psychiatry. Safe-Psych contains over 1,000 real-world psychiatric clinical notes segmented to simulate incremental evidence disclosure, with psychiatrist-derived action labels at each stage: DIAGNOSE, CLARIFY, or ABSTAIN. We evaluate multiple state-of-the-art LLMs in full-information and sequential settings. Our findings show that capability does not ensure calibration: even strong models struggle under incomplete clinical information, with under-abstention exceeding 60% for most models and safety-aware prompting reducing premature commitment only by shifting errors toward excessive abstention. In sequential evaluation, models frequently diagnose before sufficient evidence is available and rarely seek clarification unless explicitly prompted; these premature diagnoses are less accurate than on-time diagnoses. Overall, Safe-Psych reveals a limitation across the evaluated models: recognizing when clinical evidence is incomplete and additional information is needed. We release Safe-Psych to support research on improving LLM safety in healthcare.
发表机构
- National University of Science and Technology Politehnica Bucharest(布加勒斯特理工大学国立科技大学)
- Psychiatric Hospital Doctor Gheorghe Preda(格奥尔基·普雷达医生精神病医院)
- Oslo Metropolitan University(奥斯陆都市大学)
- Kristiania University of Applied Sciences(克里斯蒂亚尼亚应用科学大学)
- SimulaMet(SimulaMet公司)
- Lucian Blaga University of Sibiu(锡比乌卢西安·布拉加大学)
机构由 AI 辅助整理,请以论文原文为准。