发表机构
Soongsil University; KAIST(崇实大学; 韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SWORD基准,通过扰动Wikidata三元组生成多语言事实错误陈述,揭示LLM依赖分布熟悉度而非事实验证,并在亚洲语言上出现显著性能下降。
AI 中文摘要
现代大型语言模型(LLM)展示了令人印象深刻的多语言能力,然而标准基准主要奖励选择正确答案,而非评估真正的事实理解。我们引入了系统性的基于 Wikidata 的对象-关系失真(SWORD)基准,用于评估模型是否在不同语言中一致地拒绝事实错误。SWORD 通过对 Wikidata 三元组进行受控扰动,生成语法结构良好但事实不正确的陈述,涵盖八种广泛使用的语言,扰动范围从随机实体替换到语义上合理的基于属性的选择。我们的基于失真的评估揭示了两个被传统基准完全掩盖的关键见解。首先,模型在语义合理的失真上反而取得了更高的准确率,而在无意义的随机替换上准确率较低,这表明模型依赖于分布熟悉度而非真正的事实验证。其次,在不同语言中基线准确率相当的模型,在呈现失真陈述时,特别是在(东亚)亚洲语言上表现出显著的性能下降,某些模型的跨语言性能差距高达 28 个百分点(相对降低 49%)。这些发现表明,多语言事实推理涉及不对称的能力,而聚合准确率指标系统性地掩盖了这些能力。
英文摘要
Modern LLMs demonstrate impressive multilingual performance, yet standard benchmarks primarily reward selecting correct answers rather than evaluating genuine factual understanding. We introduce Systematic Wikidata-based Object-Relation Distortion (SWORD), a benchmark that evaluates whether models consistently reject factual errors across languages. SWORD generates syntactically well-formed but factually incorrect statements in eight widely spoken languages through controlled perturbations of Wikidata triples, ranging from random entity substitutions to semantically plausible property-based selections. Our distortion-based evaluation surfaces two critical insights that remain entirely obscured by conventional benchmarks. First, models counterintuitively achieve higher accuracy on semantically plausible distortions than on nonsensical random substitutions, suggesting reliance on distributional familiarity rather than genuine factual verification. Second, models exhibiting comparable baseline accuracy across languages show substantial performance degradation specifically on (East) Asian languages when presented with distorted statements, with cross-lingual performance gaps reaching up to 28 percentage points (49\% relative reduction) in some models. These findings demonstrate that multilingual factual reasoning involves asymmetric capabilities that aggregate accuracy metrics systematically obscure.
Comments20 pages, 12 figures, 6 tables (including appendix)