arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨语言表示对齐:语言无关空间中的令牌级最优传输

Cross-Lingual Representation Alignment by Token-Level Optimal Transport in a Language-Agnostic Space

Taisei Yamamoto, Ryoma Kumon, Danushka Bollegala, Hitomi Yanaka

arXiv 2609.06381首次发表:更新:

发表机构

The University of Tokyo; Riken; University of Liverpool; Tohoku University(东京大学; 理化学研究所; 利物浦大学; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CAROT方法,通过令牌级最优传输在语言无关空间中对齐LLM表示,保留语言特定信息,提升多语言性能并保持语言一致性,优于现有方法。

AI 中文摘要

跨语言对齐(CLA)旨在对齐大型语言模型(LLMs)在不同语言中的表示,从而实现跨语言迁移以提升多语言能力。以往的CLA方法往往忽略表示中编码的语言特定信息,仅考虑句子级对齐,这可能导致次优性能及输入输出语言不匹配的问题。我们提出了CAROT(通过最优传输在语言无关空间中实现表示的跨语言对齐),该方法包含两个步骤:识别LLMs内部状态中的语言特定表示,并通过最优传输在令牌级对齐跨语言的与语言无关的表示,同时明确保留语言特定表示。推理时引导实验表明,由CAROT计算的表示是有效的对齐目标,在保持输入输出语言一致性的同时,将多语言性能提升了最多11.2个准确率百分点。我们进一步将CAROT获得的表示作为训练目标,内化对齐后的表示。训练后的模型在18个评估设置(3个模型×3个任务×ID/OOD语言)中的11个上优于现有CLA方法。我们的工作为LLMs中CLA的有效对齐目标提供了见解。代码可在以下https URL获取。

英文摘要

Cross-lingual alignment (CLA) aims to align the representations of large language models (LLMs) across languages, enabling cross-lingual transfer to improve multilingual capabilities. Previous CLA methods often ignore language-specific information encoded in representations and only consider sentence-level alignment, which may lead to suboptimal performance and input-output language mismatch. We propose CAROT (Cross-Lingual Alignment of Representations in a Language-Agnostic Space via Optimal Transport), which consists of two steps: identifying language-specific representations in LLMs' internal states and aligning language-agnostic representations across languages at the token level by optimal transport, while explicitly preserving language-specific representations. Inference-time steering experiments show that the representations computed by CAROT are effective alignment targets, improving multilingual performance by up to 11.2 points in accuracy while maintaining input-output language consistency. We further use the representations obtained by CAROT as training targets, internalizing the aligned representations. The trained models outperform existing CLA methods in 11 of 18 evaluation settings (3 models $\times$ 3 tasks $\times$ ID/OOD languages). Our work provides insights into what constitutes effective alignment targets for CLA in LLMs. Code is available at https://github.com/ynklab/CAROT

CommentsAccepted to EMNLP 2026 main

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑