CARE-Bench:面向患者的大语言模型分诊基准测试
CARE-Bench: Benchmarking Patient-Facing LLM Triage
AI总结:
本研究推出面向患者的大语言模型分诊基准CARE-Bench,评估11个模型后发现提示可提升多数模型性能但仍存阈值错误,表明患者分诊需明确评估行动时机而非仅靠简单提示。
AI中文摘要:
面向患者的医疗大语言模型(LLM)和智能体在临床医生接触前日益频繁地回答症状相关问题,其中关键的安全问题是用户接下来应采取何种行动。我们推出CARE-Bench,这是一个基于真实来源的基准测试,将面向患者的序贯分诊评估作为每轮当前行动的四标签任务。CARE-Bench包含500个案例和1059个经重构的患者披露前缀,这些前缀源自医疗对话、咨询及后续问题来源。我们在269个保留回合中,针对未提示和极简提示的开放式协议,评估了11个模型,使用固定的GPT-5.5映射器将每个响应编码为四标签行动空间。未提示的宏F1值仍较低,范围为31.2至50.4。提示提升了11个模型中的10个,提示后的宏F1值范围为46.9至63.4,但仍存在大量阈值错误。提示后的模型常建议在获取必要澄清前就采取护理措施;当正确行动是要求更多信息时,仅33.5%的提示后输出保留了该步骤。提示后这些错误依然存在,表明面向患者的分诊并非简单的提示问题,支持在部署前对行动时机进行明确评估。
英文摘要:
Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is what action the user should take next. We introduce CARE-Bench, a source-grounded benchmark that evaluates sequential patient-facing triage as a four-label per-turn current-action task. CARE-Bench contains 500 cases and 1,059 evaluated patient-disclosure prefixes reconstructed from medical dialogue, consultation, and follow-up-question sources. We evaluate 11 models on 269 held-out rounds under unprompted and minimally prompted open-ended protocols, using a fixed GPT-5.5 mapper to code each response into the four-label action space. Unprompted macro-F1 remains low, ranging from 31.2 to 50.4. Prompting improves 10 of 11 models, with prompted macro-F1 ranging from 46.9 to 63.4, but substantial threshold errors remain. Prompted models often recommend care before needed clarification is obtained; when the correct action was to ask for more information, only 33.5% of prompted outputs preserved the step. The persistence of these errors after prompting suggests that patient-facing triage is not a simple prompting problem and supports explicit evaluation of action timing before deployment.