arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07687cs.CL

多语言LLM在医疗问题中的跨语言一致性视角

Perspectives on Cross-Lingual Consistency in LLMs for Medical Questions

Minh Duc Bui, Mario Sanz-Guerrero, Abteen Ebrahimi, Sagi Shaier, Peter Herbert Kann, Manuel Mager, Katharina von der Wense

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过调查356名专业人员,探讨多语言LLM在医疗问题中应保持一致还是适应文化,发现观点分歧,且LLM无法重现人类差异,需更多实证证据。

中文摘要 AI 辅助

多语言LLM是否应跨输入语言一致地回答医疗问题,还是应使回答适应文化线索?现有的多语言医疗基准通常假设医学上正确的答案应跨语言保持一致,并将跨语言变异视为模型错误。相反,文化适应研究认为,适当的医疗答案在不同情境下可能合法地有所不同。我们通过这两种视角回顾了多语言医疗NLP文献,识别出三个空白:利益相关者视角有限(例如,医疗专业人员的视角),缺乏关于哪种方法更能服务用户的实证证据,以及没有能够区分普遍正确与文化特定案例的基准。为解决第一个空白,我们对来自三个国家(德国、西班牙和美国)的三个利益相关者群体(医疗、NLP和人类学专业人员)的356名参与者进行了调查。人类学家一致支持适应,而医疗和NLP受访者意见分歧,美国和欧洲医疗专业人员之间存在显著差异。以职业和国家人物角色提示的LLM未能重现这种差异,高估了NLP和医疗人物角色中跨语言一致性的偏好。我们得出结论,目前既不能明确认为一致性更优,也不能明确认为适应性更优,这凸显了需要实证证据来确定哪种方法在不同文化情境中更能服务用户。

英文摘要

Should multilingual LLMs answer medical questions consistently across input languages, or adapt responses to cultural cues? Existing multilingual medical benchmarks usually assume that medically correct answers should remain consistent across languages and treat cross-lingual variation as model error. In contrast, cultural adaptation research argues that appropriate medical answers may legitimately differ across contexts. We review the multilingual medical NLP literature through these two perspectives, we identify three gaps: limited stakeholder perspectives (e.g., of medical professionals), a lack of empirical evidence on which approach better serves users, and no benchmarks capable of distinguishing universally correct from culture-specific cases. To address the first gap, we survey 356 participants across three stakeholder groups (medical, NLP, and anthropology professionals) in three countries (Germany, Spain, and the United States). Anthropologists consistently favor adaptation, while medical and NLP respondents remain divided, with notable divergence between U.S. and European medical professionals. LLMs prompted with profession and country personas fail to reproduce this variation, overestimating cross-lingual consistency preference among NLP and medical personas. We conclude that neither consistency nor adaptation can currently be considered clearly preferable, highlighting the need for empirical evidence on which approach better serves users across cultural contexts.

发表机构

  • Johannes Gutenberg University Mainz(美因茨约翰内斯·古腾堡大学)
  • University of Colorado Boulder(科罗拉多大学博尔德分校)
  • University of Marburg(马尔堡大学)
  • Universidad Iberoamericana(伊比利亚美洲大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑