在多样性、公平性与包容性提示下医学语言模型的人口统计信息注入问题
Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts
- Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究发现,在医学问题后附加DEI提示会使47个医学语言模型的人口统计信息注入率升至33.1%,部分虚构人口统计信息会误导模型输出错误答案,该效应源于公平性内容,是通用机制的体现。
AI中文摘要:
临床人工智能(Clinical-AI)领域的指导意见日益建议,通过提示语言模型使其在推理过程中关注多样性、公平性与包容性(DEI)。我们测量到一种会错误表征患者的副作用:在一个医学问题后附加一句DEI提示,会导致模型添加问题从未提及的患者人口统计属性(种族、社会经济地位、性别),本质上是改写了患者的身份,我们将此称为人口统计信息注入。在47个模型、4个医学基准及由经过验证的模型评判流程评分的376000条响应中,一条DEI提示使所有47个模型的注入率从0.7%提升至33.1%(增幅达47倍),该效应源于公平性内容而非附加长度(与长度匹配的对照组相比增幅为18倍,p值为1.4×10^-14)。大部分添加内容是不改变答案的一般人群表述,但有一小部分会将属性附加到特定患者或改变所选选项(占响应的0.25%-2.4%,其中99.8%指向错误选项),此时虚构的人口统计信息会改变模型推荐的答案。措辞会使该效应在14%至56%之间变化。DEI提示只是更通用机制的一个示例,任何推动模型推理方式的指令都可能使其添加未被要求的细节,包括关于患者的细节。本研究中将标记的输出视为模型错误,而非临床指导意见。
英文摘要:
Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI). We measure a side effect that misrepresents patients: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes (race, socioeconomic status, sex) the question never stated, in effect rewriting who the patient is. We call this demographic injection. Across 47 models, four medical benchmarks, and 376,000 responses scored by a validated model-judge pipeline, a single DEI prompt raises the injection rate from 0.7% to 33.1% (47x) in all 47 of 47 models, attributable to the equity content rather than to added length (18x above a length-matched control; p=1.4x10^-14). Most added content is a general population statement that leaves the answer unchanged, but a smaller subset attaches an attribute to the specific patient or changes the selected option (0.25-2.4% of responses, 99.8% toward the incorrect option), where the invented demographic changes the answer the model recommends. Phrasing scales the effect from 14% to 56%. DEI prompts are just one example of a more general mechanism. Any instruction that nudges how a model reasons can make it add unrequested details, including details about the patient. Flagged outputs are treated as model errors under study, not clinical guidance.