arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2409.01497cs.CL

DiversityMedQA:使用大型语言模型评估医学诊断中的人口统计学偏见

DiversityMedQA: Assessing Demographic Biases in Medical Diagnosis using Large Language Models

  • Algoverse AI Research(Algoverse人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

Rajat Rawat, Hudson McBride, Dhiyaan Nirmal, Rajarshi Ghosh, Jong Moon, Dhruv Alamuri, Sean O'Brien, Kevin Zhu

更新

AI总结:

本文提出 DiversityMedQA 基准,通过扰动 MedQA 考题并加入验证过滤策略,评估 LLM 在不同性别、族裔患者设定下的诊断表现差异,为发现和缓解人口统计学偏见提供资源。

AI中文摘要:

随着大型语言模型(LLM)在医疗保健领域得到越来越多的应用,人们对其易受人口统计学偏见影响的担忧也在不断增加。我们推出了 DiversityMedQA,这是一个新颖的基准,旨在评估 LLM 对涵盖不同患者人口统计学特征(如性别和族裔)的医学查询所作出的回答。通过对 MedQA 数据集中由医学委员会考试题目组成的问题进行扰动,我们创建了一个能够捕捉不同患者档案下医学诊断细微差异的基准。研究结果表明,在针对这些人口统计学变化进行测试时,模型性能存在显著差异。此外,为确保扰动准确,我们还提出了一种过滤策略来验证每一次扰动。通过发布 DiversityMedQA,我们为评估和缓解 LLM 医学诊断中的人口统计学偏见提供了一种资源。

英文摘要:

As large language models (LLMs) gain traction in healthcare, concerns about their susceptibility to demographic biases are growing. We introduce {DiversityMedQA}, a novel benchmark designed to assess LLM responses to medical queries across diverse patient demographics, such as gender and ethnicity. By perturbing questions from the MedQA dataset, which comprises medical board exam questions, we created a benchmark that captures the nuanced differences in medical diagnosis across varying patient profiles. Our findings reveal notable discrepancies in model performance when tested against these demographic variations. Furthermore, to ensure the perturbations were accurate, we also propose a filtering strategy that validates each perturbation. By releasing DiversityMedQA, we provide a resource for evaluating and mitigating demographic bias in LLM medical diagnoses.

补充信息

↑