arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.00724cs.CL

MSQA:一个原生来源的多语言多文化SimpleQA基准

MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

  • M-A-P
  • ByteDance Seed(字节跳动Seed)
  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

Xianru Chen, Yukai Huang, Mingxiang Chen, Xinping Lei, Fangbing Deng, Jin Chen, Ge Zhang, Wenhao Huang, Jiaheng Liu

AI总结:

提出MSQA基准,包含1064个原生问题覆盖11种语言和5个文化维度,评估18个LLM发现文化能力随预训练暴露程度下降,且推理时补救措施无效。

AI中文摘要:

多语言流利性常常引发一个更强的假设:一个能说用户语言的模型也必须理解该语言编码的文化。我们称之为文化对齐的幻觉。为了直接检验这一假设,我们引入了MSQA,一个包含1064个原生来源问题的基准,涵盖11个语言组、五个文化维度和三个难度层级。与翻译基准不同,MSQA针对本地化知识,并减少了来自以英语为中心的跨语言迁移的捷径。评估18个LLM,我们发现显著的文化退化以及明显的局部性效应:文化能力更紧密地追踪预训练暴露程度,而非一般推理能力。我们进一步表明,常见的推理时补救措施并不能消除这种幻觉。模型对不熟悉的文化问题仍然过度自信,重复采样产生不稳定而非可靠的正确性,检索增强对长尾事实的帮助不均匀。这些发现表明,文化对齐不能仅从多语言能力推断,并且需要比推理时的校准、采样或检索更深入的干预。

英文摘要:

Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that language. We call this the Illusion of Cultural Alignment. To test this assumption directly, we introduce MSQA, a benchmark of 1,064 natively sourced questions across 11 language groups, five cultural dimensions, and three difficulty tiers. Unlike translated benchmarks, MSQA targets locally grounded knowledge and reduces shortcuts from English-centric cross-lingual transfer. Evaluating 18 LLMs, we find substantial cultural degradation and a pronounced Locality Effect: cultural competence tracks pre-training exposure more closely than general reasoning ability. We further show that common inference-time remedies do not dissolve the illusion. Models remain overconfident on unfamiliar cultural questions, repeated sampling yields unstable rather than reliable correctness, and retrieval augmentation helps unevenly on long-tail facts. These findings indicate that cultural alignment cannot be inferred from multilingual ability alone and requires deeper intervention than calibration, sampling, or retrieval at inference time

补充信息

↑