arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03367cs.AIcs.CL

多语言GSM-Symbolic:什么决定了跨语言的能力迁移?

Multilingual GSM-Symbolic: What determines capability transfer across languages?

  • Aarhus University(奥胡斯大学)
  • Danish Foundation Models(丹麦基础模型公司)
  • University of Alabama(阿拉巴马大学)
  • Alexandra Institute(亚历山德拉研究所)
  • Indian Institute of Technology Kharagpur(印度理工学院卡拉格普尔分校)
  • The University of Tokyo(东京大学)
  • IT University of Copenhagen(哥本哈根信息技术大学)
  • Massachusetts General Hospital(马萨诸塞州总医院)
  • Bocconi University(博科尼大学)
  • University of Southern Denmark(南丹麦大学)
  • University of Iceland(冰岛大学)
  • University of the Faroe Islands(法罗群岛大学)
  • Zendesk(Zendesk公司)
  • University of Copenhagen(哥本哈根大学)
  • Indian Institute of Technology Madras(印度理工学院马德拉斯分校)
  • National Library of Sweden(瑞典国家图书馆)

机构由 AI 辅助整理,请以论文原文为准。

Kenneth Enevoldsen, Riley Herchert, Sofie Mosegaard, Dan Saattrup Smart, Simon Enni, Isaac Chung, Sofie Bruun, Ayush Sunil Munot, Max Müller-Eberstein, Adnan El… 展开作者

Kenneth Enevoldsen, Riley Herchert, Sofie Mosegaard, Dan Saattrup Smart, Simon Enni, Isaac Chung, Sofie Bruun, Ayush Sunil Munot, Max Müller-Eberstein, Adnan El-Assadi, Elisa Bassignana, Gianluca Barmina, Hafsteinn Einarsson, Iben Nyholm Debess, Linda Freienthal, Lukas Galke Poech, Mike Zhang, Nicolas Legrand, Vladimir Salnikov, Yevhen Kostiuk, Zafar Hussain, Sagandeep Kaur, Agnes Toftgård, Marie Mattson, Kristoffer Nielbo

AI总结:

本研究提出多语言GSM-Symbolic数据集,量化模型大小、语言资源、推理和类型学距离对跨语言能力迁移的影响,并揭示模型大小和推理可缩小语言间性能差距。

AI中文摘要:

我们对于以一种语言获得的能力如何迁移到另一种语言,以及什么因素支配这种迁移知之甚少:评估依赖于不可比较的、容易饱和的数据集,并且很少联合检验其决定因素。识别什么能预测迁移将使我们能够避免对所有语言对进行详尽评估,并让开发者针对限制低资源语言性能的因素进行改进。为了评估跨语言能力迁移,我们引入了多语言GSM-Symbolic,一个可扩展的多语言数学数据集,涵盖30,000个题目匹配的问答对,跨越15种语言。它利用符号模板来防止过拟合,并通过允许从单个样本生成数百万个高质量变体来确保泛化。使用多语言GSM-Symbolic,我们量化了能力的主要决定因素:模型大小(β=1.77)、语言资源水平(β=0.77)、推理(β=0.67)和类型学距离(β=-0.25)。这种联合估计使得这些决定因素可以相互表达:在马拉地语中评估的32B模型表现得像英语中的10B模型。我们的发现对模型开发者具有重要意义,表明模型大小和推理缩小了低资源和高资源语言之间的性能差距(分别为β=-0.27和β=-0.20),而类似的杠杆对类型学上遥远的语言几乎没有影响。总体而言,我们的分析框架解释了92%的语言间变异,但仅解释了23%的模型-语言变异,并在6.0个百分点(r=.96)内预测模型在未见语言上的性能。结合目标语言中仅10个模板的测量,可将此降低到4.19个百分点,从而在几乎没有或没有下游数据集的情况下实现合理的性能估计。

英文摘要:

We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluation across all language pairs and let developers target the factors that limit performance in low-resource languages. To evaluate cross-lingual capability transfer, we introduce Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages. It utilises symbolic templates to prevent overfitting and ensure generalisation by allowing generation of millions of high-quality variations from a single sample. Using Multilingual GSM-Symbolic, we quantify the largest determinants of capability as model size ($β= 1.77$), language resource level ($β= 0.77$), reasoning ($β= 0.67$) and typological distance ($β= -0.25$). This joint estimation allows these determinants to be expressed in terms of one another: a 32B model evaluated in Marathi performs like a 10B model in English. Our findings have important implications for model developers, showing that model size and reasoning narrow the performance gap between low- and high-resource languages ($β= -0.27$ and $β= -0.20$, respectively), while similar levers have little or no effect on typologically distant languages. Overall, our analysis framework explains 92% of between-language variation, but only 23% of the model-by-language variation, and predicts a model's performance on an unseen language within 6.0pp (r=.96). Incorporating measurements from just 10 templates in the target language reduces this to 4.19pp, enabling reasonable estimates of performance with little or no downstream dataset.

↑