arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多语言语言模型对齐的元学习偏好

Meta-Learning Preferences for Multilingual LLM Alignment

Jiaying Lin, Seongho Son, Nam Phuong Tran, Long Tran-thanh, Ilija Bogunovic, Debmalya Mandal

arXiv 2607.13315首次发表:更新:

发表机构

University of Warwick; University College London; University of Basel(华威大学; 伦敦大学学院; 巴塞尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多语言环境下低资源语言对齐数据不足问题,提出元学习框架,利用其他语言偏好数据学习可转移初始化,经理论和实验验证,该方法在极低资源设置下比基线方法胜率最多提高28%,且在多场景表现优异。

AI 中文摘要

跨语言的人类偏好数据可用性不均,给多语言环境下的大语言模型对齐带来重大挑战。为解决低资源语言对齐中数据不足的问题,我们提出了一个用于从人类反馈强化学习和直接偏好优化的元学习框架。通过利用其他语言的偏好数据,该框架学习可转移的初始化,以最少的数据实现对目标语言的有效适应。我们为元奖励建模和元策略优化设置提供了理论保证,并通过实验证明了该方法在多语言基准测试中的有效性。在仅有100个目标语言偏好样本的极低资源设置中,我们的方法比基线方法的胜率提高了28%,并在多种目标语言和模型规模上持续优于基线。我们的方法在元训练语言的不同组合以及与目标语言不同的语言距离下都保持了这些优势。

英文摘要

Unequal availability of human preference data across languages poses a significant challenge for aligning large language models in multilingual settings. To address the lack of sufficient data in low-resource language alignment, we propose a meta-learning framework for Reinforcement Learning from Human Feedback and Direct Preference Optimization. By leveraging preference data from other languages, our framework learns a transferable initialization that enables effective adaptation to a target language with minimal data. We provide theoretical guarantees for both the meta-reward modeling and meta-policy optimization settings, and empirically demonstrate the effectiveness of our approach on multilingual benchmarks. In an extremely low-resource setting with only 100 target-language preference samples, our approach achieves up to $28\%$ win-rate improvements over baseline methods, and consistently outperforms baselines across multiple target languages and model scales. Our approaches retain these advantages across different combinations of meta-training languages and varying linguistic distances from the target languages.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑