AI 中文总结
研究多轮医疗对话中模型对误解的处理,引入ThReadMed-QA数据集,用基于准则的框架评估五个大语言模型,发现前沿模型在后续轮次表现大幅下降,错误传播致性能降,强调需捕捉多轮行为的评估框架。
AI 中文摘要
寻求医疗信息的患者常提出包含错误假设或误解的问题。安全的医疗沟通不仅要回答问题,还需识别并纠正潜在错误信念。这种互动自然展开于多轮,如今与大语言模型的互动也如此。但当前评估框架未涵盖误解在对话中出现、持续或演变的情况。为研究此,我们引入ThReadMed-QA,一个包含2437个医患对话线程、8204个问答对的多轮医疗对话数据集。我们用基于准则的大语言模型评判框架评估五个大语言模型,发现即使能在单次互动中处理误解的前沿模型,在后续轮次中表现也大幅下降。神谕分析表明,性能下降多由错误传播导致,即使在正确语境下性能也不完美。这凸显了捕捉多轮行为的评估框架的必要性。
英文摘要
Patients seeking medical information often ask questions that embed incorrect assumptions or misconceptions. In such cases, safe medical communication requires not only answering the question, but identifying and correcting the underlying false belief. These interactions naturally unfold over multiple turns, a pattern now mirrored in interactions with LLMs. Yet current evaluation frameworks do not capture model behavior in these settings, where misconceptions can emerge, persist, or evolve over the course of a conversation. Whether LLMs can reliably correct such misconceptions over time remains largely unexamined. To study this, we introduce ThReadMed-QA, a multi-turn medical dialogue dataset of 2,437 patient-physician conversation threads comprising 8,204 question-answer pairs, derived from real patient interactions on AskDocs. This dataset enables systematic evaluation of whether models can detect and correct misconceptions under a multi-turn context. We evaluate five LLMs using a rubric-based LLM-as-a-Judge framework that scores responses based on their ability to identify and correct misconceptions. Our experiments reveal a consistent pattern: even frontier models that can address misconceptions in a single interaction degrade substantially over subsequent turns. GPT-5 and Claude-Haiku correct these false presuppositions around 85% on initial questions but drop to roughly 50% within two follow-ups. An oracle analysis replacing prior model outputs with physician responses shows that much of the degradation is driven by error propagation, while performance remains imperfect even under correct context. Even when models tend to correct misconceptions initially, their performance degrades substantially over later turns, leading to inconsistent and potentially unsafe guidance in patient-facing settings and highlighting the need for evaluation frameworks that capture multi-turn behavior.
CommentsAccepted to MLHC 2026