发表机构
Eastern University(东方大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究定义跨语言理解差距(CLCG),通过ParallelQA-18评估5种模型,发现英语能力无法同等迁移,低资源语言用户的模型质量可能被英语中心评估高估。
AI 中文摘要
人们在评估语言模型时,通常认为模型在英语中展现的能力,在以其他语言呈现相同内容时也同样可用。传统多语言基准很少能在保持内容、问题、参考答案、模型和评估单元一致的情况下,单独分离出语言变量。我们将跨语言理解差距(Cross-Lingual Comprehension Gap, CLCG)定义为:当相同的内容和问题以目标语言而非英语呈现时,模型响应质量的下降幅度。我们使用由专业人工翻译的平行语料库ParallelQA-18,对来自5个实验室的5种模型,在覆盖18种语言的150篇文章分层样本上开展评估(英语作为参考;葡萄牙语作为高资源基准;16种目标语言涵盖Joshi等人2020年分类的0-4类)。采用项内设计,仅改变段落语言。主要估计量对比英语与合并目标语言在较高复杂度开放式问题上的Token-F1微观均值,并计算文章聚类自举区间。合并后的主要CLCG为0.078(95%置信区间0.072-0.084),相对于英语得分下降约17%;同语言宏观摘要为0.077。扣除葡萄牙语后,宏观差距为0.016(95%置信区间0.013-0.020)。语言层面的CLCG与Joshi资源类呈负相关(rho=-0.594,p=0.015,n=16)。在盲法配对人工评估中,高资源语言对应的响应在61.6%的决定性判断中更受偏好(估计偏好概率0.655,95%置信区间0.558-0.741)。研究表明,不能假设英语中展现的能力会同等迁移到其他语言;以英语为中心的评估可能会高估低资源语言用户的模型质量。
英文摘要
Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other languages. Traditional multilingual benchmarks rarely isolate language while holding content, question, reference answer, model, and evaluation unit constant. We define the Cross-Lingual Comprehension Gap (CLCG) as the reduction in response quality when the same content and question are presented in a target language rather than in English. Using ParallelQA-18, a professionally human-translated parallel corpus, we evaluate five models from five laboratories on a stratified sample of 150 articles across 18 languages (English reference; Portuguese high-resource baseline; 16 targets spanning Joshi et al. 2020 classes 0-4). A within-item design varies only passage language. The primary estimator contrasts English versus pooled target-language Token-F1 micro-means on higher-complexity open-ended questions, with article-cluster bootstrap intervals. The primary pooled CLCG is 0.078 (95% CI 0.072-0.084), about a 17% reduction relative to the English score; the equal-language macro summary is 0.077. Net of Portuguese, the macro gap is 0.016 (95% CI 0.013-0.020). Language-level CLCG is negatively associated with Joshi resource class (rho = -0.594, p = 0.015, n = 16). In blinded paired human evaluations, higher-resource responses are preferred in 61.6% of decisive judgments (estimated preference probability 0.655, 95% CI 0.558-0.741). Capabilities shown in English should not be assumed to transfer equally to other languages; English-centered evaluations may overestimate quality for users of low-resource languages.
Comments55 pages, 17 figures. Submitted to Computational Linguistics (MIT Press / ACL). Supplementary Material: 55 pages, 4 figures