arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01921cs.CLcs.AI

跨语言对齐用于使用MoE路由器的仅解码器模型

Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal, Nanyun Peng

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出利用MoE路由器输出作为对齐目标,在仅解码器模型中实现跨语言对比学习,实验证明该方法能提升多语言性能。

中文摘要 AI 辅助

跨语言对比学习一直是多语言编码器训练的核心组成部分,但由于多语言分词的不同,在仅解码器的大型语言模型(LLM)中无法显式对齐表示。然而,越来越多的研究表明,即使在LLM中,更高的跨语言表示对齐也能带来更好的跨语言迁移。在本文中,我们提出了一种新颖的方法,根据现代LLM的架构约束重新构想跨语言对比学习。我们不是对隐藏状态应用辅助对齐损失,而是建议使用混合专家(MoE)路由器的输出作为对齐目标。路由器输出更适合对多个标记进行池化,从而在序列级别实现更可靠的跨语言比较。在四个开源MoE上进行的受控持续预训练实验表明,加入这种路由损失也能跨语言对齐底层隐藏表示。最重要的是,这种损失提高了我们在多样化评估套件上的多语言性能,展示了跨语言MoE路由器对齐的潜力。

英文摘要

Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine cross-lingual contrastive learning given the architectural constraints of modern LLMs. Rather than applying an auxiliary alignment loss on hidden states, we propose using the outputs of the mixture-of-experts (MoE) routers as the target for alignment. Router outputs lend themselves better to pooling over many tokens, enabling more reliable cross-lingual comparisons at the sequence-level. Controlled continual pre-training experiments on four open-source MoEs show that incorporating this routing loss also aligns the underlying hidden representations across languages. Most importantly, this loss improves multilingual performance on our diverse evaluation suite, demonstrating the potential of cross-lingual MoE router alignment.

发表机构

  • University of California, Los Angeles(加州大学洛杉矶分校)
  • Haverford College(哈弗福德学院)
  • MILA - Quebec AI Institute & McGill University(MILA - 魁北克人工智能研究所 & 麦吉尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑