arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13036cs.CLcs.AI

诊断前先询问:Safe-Psych,一种用于精神病学领域大语言模型的顺序评估基准

Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

Oriana Presacan, Andreea Grama, Larisa Irimină, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. Băcilă, Bogdan Ionescu, Michael A. Riegler

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对大语言模型在精神病学决策支持中处理证据不完整问题,引入顺序评估基准Safe-Psych,通过模拟真实临床记录及精神科医生行动标签,评估多个模型,发现模型存在过早诊断等问题,揭示其局限性,为改善模型安全性研究提供支持。

中文摘要 AI 辅助

大语言模型(LLMs)越来越多地用于医疗保健决策支持,但临床证据往往不完整或不断变化。当可用信息不足以支持可靠答案时,模型应要求澄清或弃权,而不是提供无根据的回答。然而,现有的医学基准通常假设一开始就有完整信息。我们引入了Safe-Psych,这是一个用于评估LLMs如何处理临床精神病学中不断变化的诊断不确定性的顺序基准。Safe-Psych包含1000多个真实世界的精神病临床记录,这些记录被分割以模拟渐进式证据披露,每个阶段都有精神科医生给出的行动标签:诊断、澄清或弃权。我们在全信息和顺序设置中评估了多个先进的LLMs。研究结果表明,能力并不能确保校准:即使是强大的模型在临床信息不完整的情况下也会挣扎,大多数模型的弃权不足超过60%,安全意识提示只是将错误转向过度弃权,从而减少过早承诺。在顺序评估中,模型经常在没有足够证据的情况下进行诊断,除非明确提示,否则很少寻求澄清;这些过早诊断的准确性低于及时诊断。总体而言,Safe-Psych揭示了所评估模型的一个局限性:识别临床证据何时不完整以及何时需要额外信息。我们发布Safe-Psych以支持改善医疗保健中LLM安全性的研究。

英文摘要

Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving. When the available information is insufficient to support a reliable answer, models should request clarification or abstain rather than provide unsupported responses. Existing medical benchmarks, however, typically assume that complete information is available upfront. We introduce Safe-Psych, a sequential benchmark for evaluating how LLMs handle evolving diagnostic uncertainty in clinical psychiatry. Safe-Psych contains over 1,000 real-world psychiatric clinical notes segmented to simulate incremental evidence disclosure, with psychiatrist-derived action labels at each stage: DIAGNOSE, CLARIFY, or ABSTAIN. We evaluate multiple state-of-the-art LLMs in full-information and sequential settings. Our findings show that capability does not ensure calibration: even strong models struggle under incomplete clinical information, with under-abstention exceeding 60% for most models and safety-aware prompting reducing premature commitment only by shifting errors toward excessive abstention. In sequential evaluation, models frequently diagnose before sufficient evidence is available and rarely seek clarification unless explicitly prompted; these premature diagnoses are less accurate than on-time diagnoses. Overall, Safe-Psych reveals a limitation across the evaluated models: recognizing when clinical evidence is incomplete and additional information is needed. We release Safe-Psych to support research on improving LLM safety in healthcare.

发表机构

  • National University of Science and Technology Politehnica Bucharest(布加勒斯特理工大学国立科技大学)
  • Psychiatric Hospital Doctor Gheorghe Preda(格奥尔基·普雷达医生精神病医院)
  • Oslo Metropolitan University(奥斯陆都市大学)
  • Kristiania University of Applied Sciences(克里斯蒂亚尼亚应用科学大学)
  • SimulaMet(SimulaMet公司)
  • Lucian Blaga University of Sibiu(锡比乌卢西安·布拉加大学)

机构由 AI 辅助整理,请以论文原文为准。

↑