arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HealMed:面向医学领域的大语言模型多语言评估

HealMed: Multilingual Evaluation of Large Language Models in Medicine

Yingjian Chen, Fan Gao, Sherry T. Tong, Haoyu Zhang, Aosong Feng, Kevin W. Jin, Xing Wu, Jinghui Lu, Abdul Samad, Akbar Faruqi, Cesar Caraballo, Cibele Brandão, Dhruva, Gupta, Eunji Jeon, Gabriel Madera-Santiago, Geon Lee, Hugo Toshio Itikawa, Insook Cho, Isabelli Martins, Isarar Siddique, Israr Ahmed, Jihyo Kwak, Kanyakorn Veerakanjana, Luis Guilherme Cardoso, Minjin Kim, Piyalitt Ittichaiwong, Renee Dua, Santiago Gudiño-Rosales, Xiujie Chen, Zeo Lapalus, Zixin Xu, Michihiro Yasunaga, Rex Ying, Heuiseok Lim, Jaewoo Kang, Chanjun Park, Hang Jiang, Ethan Goh, Hyunjae Kim, Edison Marrese-Taylor, Yusuke Iwasawa, Yutaka Matsuo, Qingyu Chen, Irene Li

arXiv 2608.19981首次发表:更新:

发表机构

HealMed Research Team(HealMed研究团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出医学领域大语言模型多语言评估基准HealMed,经多国专家历时两年开发,实验发现低资源语言性能下降明显,专有模型更稳定,医学专业知识不保证多语言鲁棒性,翻译质量影响评估结果。

AI 中文摘要

我们提出HealMed——一个经专家审核的医学领域大语言模型多语言评估基准。HealMed包含9种语言各1000个样本,源自9个数据集,涵盖三种任务形式:多项选择题问答(MCQA)、自然语言推理(NLI)和开放式问答。该基准由来自9个国家和地区的23名医师及医学专家历时两年开发完成,每一次翻译都由两名精通英语与对应目标语言的专家进行评估和修订。在HealMed上,低资源语言的性能下降最为明显,尽管不同语言和模型之间的差距大小存在显著差异。最强的专有模型在不同语言间表现最为稳定,而许多开源模型和医学专用模型则呈现出更大且更不稳定的差距。仅具备医学专业知识并不能确保多语言鲁棒性。此外,专家修订可能会提高或降低测得的性能,这表明翻译质量对跨语言评估结果有实质性影响。

英文摘要

We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions. Each translation was evaluated and revised by two experts fluent in English and the corresponding target language. On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models. The strongest proprietary models were the most stable across languages, whereas many open-source and medically specialized models showed larger and less consistent gaps. Medical specialization alone did not ensure multilingual robustness. Furthermore, expert revision could either raise or lower measured performance, indicating that translation quality materially affects cross-language evaluation results.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑