arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向可信临床决策支持的、关注社会决定健康因素(SDoH)的医学大语言模型叙事锚定偏差

SDoH-Aware Narrative Anchoring Bias in Medical LLMs for Trustworthy Clinical Decision Support

Ahnaf Atef Choudhury, Ramkrishna Saha

arXiv 2608.22802首次发表:更新:

发表机构

George Mason University; The University of Texas at Dallas(乔治梅森大学; 德克萨斯大学达拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对医学大语言模型的叙事锚定偏差问题,以Qwen2.5系列模型为对象开展实验,发现7B模型仍存在较高叙事敏感性误差,提出需结合正确率与叙事稳定性评估临床决策支持模型。

AI 中文摘要

医学大语言模型的评估通常以其回答临床问题的正确率为依据,这一视角虽有价值,但忽略了一项实际风险:模型即便知晓正确答案,当同一病例以不同患者视角表述时,仍可能改变自身回应。本文将该风险评估为关注社会决定健康因素(SDoH)的叙事锚定偏差。研究采用NarrativeShield SDoH MedQA,这是一个反事实医学问答数据集,每个病例以基于角色的叙事呈现,而答案键保持固定;该数据集从宽格式重塑为按病例分组的角色行。我们评估了Qwen2.5系列的三个开源指令调优大语言模型:1.5B、3B和7B。最终实验使用300个临床病例,在三种提示条件下生成8100条模型回应,报告了角色层面准确率、反事实一致性、正确一致性和叙事敏感性误差。Qwen2.5 7B的准确率达56.33%,为最佳水平,正确一致性达40.33%,同样最优;配对McNemar精确检验显示,在所有提示设置下,7B的准确率均显著高于3B。即便如此,叙事敏感性依然存在,最低误差仍达31.67%。这些结果表明,可信临床决策支持应同时通过平均正确率和医学等效患者叙事间的稳定性来评估。

英文摘要

Medical large language models are often judged by how many clinical questions they answer correctly. That view is useful, but it misses a practical risk. A model may know the right answer and still change its response when the same case is written in a different patient voice. This paper evaluates that risk as SDoH aware narrative anchoring bias. We use NarrativeShield SDoH MedQA, a counterfactual medical question answering dataset in which each case appears in persona based narratives while the answer key remains fixed. The dataset is reshaped from wide format into case grouped persona rows. We evaluate three open source instruction tuned LLMs from the Qwen2.5 family: 1.5B, 3B, and 7B. The final experiment uses 300 clinical cases and produces 8,100 model responses across three prompting conditions. We report persona level accuracy, counterfactual consistency, correct consistency, and narrative sensitivity error. Qwen2.5 7B achieves the best accuracy at 56.33 percent and the best correct consistency at 40.33 percent. Paired McNemar exact tests show significant accuracy gains for 7B over 3B in all prompt settings. Even so, narrative sensitivity remains, with the lowest error still at 31.67 percent. These results suggest that trustworthy clinical decision support should be evaluated by both average correctness and stability across medically equivalent patient narratives.

CommentsAccepted for publication at 10th International Artificial Intelligence and Data Processing Symposium (IDAP'26)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑