arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在线语言自适应采样以实现更好的分布式跨语言增益

Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains

Quang Phuoc Nguyen, Félix Gaschi, David Anugraha, Santiago Martínez Novoa, En-Shiun Annie Lee

arXiv 2609.14969首次发表:更新:

发表机构

Ontario Tech University; Doctrine; Stanford University; University of the Andes; University of Toronto(安大略理工大学; Doctrine公司; 斯坦福大学; 安第斯大学; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多语言模型跨语言对齐中均匀采样次优的问题,提出在线自适应采样策略,按损失动态调整各语言采样概率,提升极低资源语言性能并均衡增益。

AI 中文摘要

重新对齐是提高多语言语言模型跨语言迁移能力的一种有前景的方法,尤其对于极低资源语言(LRLs)而言。然而,现有的重新对齐方法依赖于跨语言的均匀随机平行句采样,这在有限的批次大小下可能不是最优的。在实践中,模型可能受益于更频繁地看到某些语言,特别是那些对齐不佳的语言,并且最优分布可以在训练过程中演变。在这项工作中,我们提出了一种简单而有效的自适应采样策略,为每种语言分配可训练的采样概率。对重新对齐损失贡献更大的语言在后续批次中被更频繁地采样,并且最优分布可以在训练过程中演变。我们的方法采用了一个内-外优化循环,开销较小,从而带来一致的性能提升,更重要的是,将增益分布到各语言上。我们观察到,与均匀重新对齐相比,使用XLM-R在所有任务上平均性能提升+0.67,使用Gemma 2 9B提升+0.60。此外,我们的方法在不同模型上具有鲁棒性。代码可在该https URL获取。

英文摘要

Realignment is a promising approach for improving the cross-lingual transfer ability of multilingual language models, particularly for extremely low-resource languages (LRLs). However, existing realignment methods rely on uniform and random sampling of parallel sentences across languages, which may be suboptimal under limited batch sizes. In practice, models may benefit from seeing certain languages more frequently, especially those that are poorly aligned, and the optimal distribution can evolve throughout training. In this work, we propose a simple yet effective adaptive sampling strategy that assigns trainable sampling probabilities to each language. Languages that contribute more to the realignment loss are sampled more frequently in subsequent batches, and the optimal distribution can evolve throughout training. Our method employs an inner-outer optimization loop with a small overhead, leading to consistent performance improvements and, more importantly, distributing the gains across languages. We observed a $+0.67$ average performance increase on all tasks with XLM-R, and $+0.60$ with Gemma 2 9B compared with uniform realignment. Furthermore, our method is robust across different models. Code available at https://github.com/felixgaschi/multilingual-alignment-and-transfer.

CommentsAccepted to ENMLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑