arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于医疗咨询的大语言模型评估时机过晚:预 formulation 差距

LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

Yining Hua, Cyrus Ayubcha, Hongbin Na, Levi Lian, Alon Gorenshtein, Yiftach Barash, Eyal Klang

arXiv 2608.17330首次发表:更新:

发表机构

Harvard T.H. Chan School of Public Health; Harvard Medical School; Australian Artificial Intelligence Institute, University of Technology Sydney; Raycaster; Beth Israel Deaconess Medical Center(哈佛陈曾熙公共卫生学院; 哈佛医学院; 悉尼科技大学澳大利亚人工智能研究所; Raycaster; 贝斯以色列女执事医疗中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究指出医疗咨询LLMs评估时机过晚的预 formulation 差距,通过实验发现指导条件可调整咨询流程与记录,建议直接以首次接触行为评估该差距。

AI 中文摘要

用于医疗咨询的大语言模型(LLMs)常是在临床问题已明确后才接受评估,而实际咨询可能始于模糊、简化或框架不当的诉求。我们在基线和就诊入门指导条件下,针对4个由医生撰写的多回合 vignette 评估了3个 API 模型,得到24个固定脚本 transcript;2个案例还采用自适应标准化患者模拟,得到12个 transcript。在基线案例-模型单元的12组中,有9组在患者未作答前就给出自我护理或家庭管理建议,而在指导条件的12组中为0;结构化交接摘要在基线的12组中为0,在指导条件的12组中为10。指导条件改变了咨询流程和记录方式,但未可靠确保引出关键事实。因此,预 formulation 差距应通过可观察的首次接触行为直接评估,而非从诊断准确性或最终答案质量推断。

英文摘要

Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24 fixed-script transcripts; two cases also used adaptive standardized-patient simulation, yielding 12 transcripts. Self-care or home-management advice before any patient answer appeared in 9 of 12 baseline case-model cells and 0 of 12 instruction cells, while structured handoff summaries appeared in 0 of 12 and 10 of 12 cells, respectively. The instruction changed sequencing and documentation, although it did not reliably ensure elicitation of decisive facts. The preformulation gap should therefore be evaluated directly through observable first-contact behavior rather than inferred from diagnostic accuracy or final-answer quality.

Comments17 pages, 3 tables. Code, cases, prompts, complete transcripts, and results: https://github.com/ningkko/preformulation-gap

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑