arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2511.01360cs.CL

翻译后更安全?印度语系中的预设鲁棒性

Safer in Translation? Presupposition Robustness in Indic Languages

  • University of Maryland, College Park (UMD) College Park, MD, USA(马里兰大学 College Park 分校)

机构由 AI 辅助整理,请以论文原文为准。

Aadi Palnitkar, Arjun Suresh, Rishi Rajesh, Puneet Puli

更新

中文总结 AI 辅助

本文构建了Cancer-Myth-Indic基准,将Cancer-Myth翻译为五种印度语系语言以保留错误预设,并在此多语言预设压力下评估了多种流行LLMs的医疗咨询表现。

中文摘要 AI 辅助

越来越多人开始向大型语言模型(LLMs)寻求医疗建议和咨询,因此评估LLMs对此类查询响应的有效性和准确性变得至关重要。尽管已有医学基准文献试图完成这一任务,但这些基准几乎全部使用英语,这导致现有文献在多语言LLM评估方面存在显著空白。在这项工作中,我们希望通过Cancer-Myth-Indic来帮助填补这一空白,这是一个印度语系基准,通过将Cancer-Myth的500项子集均匀采样自其原始类别,翻译成次大陆五种服务不足但广泛使用的语言(每种语言500项;共计2500个翻译项目)而构建。母语翻译者遵循风格指南,在翻译中保留隐含预设;这些项目包含与癌症相关的错误预设。我们在此预设压力下评估了几种流行的LLMs。

英文摘要

Increasingly, more and more people are turning to large language models (LLMs) for healthcare advice and consultation, making it important to gauge the efficacy and accuracy of the responses of LLMs to such queries. While there are pre-existing medical benchmarks literature which seeks to accomplish this very task, these benchmarks are almost universally in English, which has led to a notable gap in existing literature pertaining to multilingual LLM evaluation. Within this work, we seek to aid in addressing this gap with Cancer-Myth-Indic, an Indic language benchmark built by translating a 500-item subset of Cancer-Myth, sampled evenly across its original categories, into five under-served but widely used languages from the subcontinent (500 per language; 2,500 translated items total). Native-speaker translators followed a style guide for preserving implicit presuppositions in translation; items feature false presuppositions relating to cancer. We evaluate several popular LLMs under this presupposition stress.

补充信息

↑