相同事实,不同诊断:临床语言模型中叙事锚定的测量与缓解
Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models
浏览论文内容
中文总结 AI 辅助
本研究针对临床语言模型的叙事锚定偏差,构建含1000个USMLE案例的基准,提出NarrativeShield三智能体流水线,可大幅降低叙事锚定差距并减少不稳定决策,同时发布数据集供相关研究使用。
中文摘要 AI 辅助
用于临床诊断推理的大语言模型不仅对临床内容敏感,还会受社会语域影响。我们将这种失效模式称为叙事锚定:相同的临床事实以不同语域表达时,会导致诊断输出产生差异。与之前操纵种族、收入等明确身份标记的人口统计学偏差研究不同,我们的基准将语域作为唯一的变化渠道,不包含任何形式的人口统计学标记。我们构建了包含1000个美国医师执照考试(USMLE)临床案例的数据集,每个案例在独立审计的事实保留保证下,被重写为三个具有不同社会语域的角色,且由一个从未见过生成提示的独立模型进行验证。在覆盖三个架构家族和不同规模的七个语言模型中,直接提示下的叙事锚定在所有测试模型中均具有统计学显著性,叙事锚定差距为0.064至0.151。思维链推理和明确的去偏差指令仅能部分减少偏差,且其表面增益常因准确率下降而混淆。我们引入了NarrativeShield,一种三智能体流水线,在诊断推理开始前结构化提取并验证临床事实,将叙事锚定差距降至接近零(-0.004至0.037),并在所有模型的所有方法中实现了最低的严重不稳定决策率(DSS < 0.8),同时对大多数模型仅产生适度且符合机制预期的准确率损失。使用非指令微调的基础模型进行的压力测试表明,执行去偏差干预完全由零样本指令跟随能力决定,而非仅由提示内容决定。我们发布了经人工验证事实保留的数据集,作为研究基于语域的临床偏差的独立资源。
英文摘要
Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure mode Narrative Anchoring: identical clinical facts expressed in different registers cause diagnostic outputs to diverge. Unlike prior demographic-bias work, which manipulates explicit identity tokens such as race or income, our benchmark isolates register as the sole channel of variation, with no demographic marker present in any form. We construct a dataset of 1,000 USMLE clinical vignettes, each rewritten into three sociolinguistically distinct personas under an independently audited fact-preservation guarantee, verified by a separate model that never sees the generation prompt. Across seven language models spanning three architecture families and scales, Narrative Anchoring is statistically significant under direct prompting in every model tested, with a Narrative Anchoring Gap of 0.064 to 0.151. Chain-of-thought reasoning and explicit debiasing instructions reduce the bias only partially, and their apparent gains are frequently confounded by accuracy collapse. We introduce NarrativeShield, a three-agent pipeline that structurally extracts and verifies clinical facts before diagnostic reasoning begins, reducing the Narrative Anchoring Gap to near-zero ($-0.004$ to $0.037$) and achieving the lowest rate of severely unstable decisions (DSS $<$ 0.8) of any method across all models, at a modest and mechanistically expected accuracy cost for most models. A stress test using a non-instruction-tuned base model shows that executing a debiasing intervention at all is gated by zero-shot instruction-following ability, not prompt content alone. We release our dataset, human-validated for fact preservation, as a standalone resource for studying register-based clinical bias.
发表机构
- University of Texas at Austin(德克萨斯大学奥斯汀分校)
- Queen’s University Belfast(贝尔法斯特女王大学)
- Middlesex University London(伦敦密德萨斯大学)
机构由 AI 辅助整理,请以论文原文为准。