AI 中文总结
该研究对比分析27种跨语言对齐分数变体,提出基于PMI的翻译指标,发现用英语的跨语言对齐可相当或更好预测LLMs翻译性能,为LLMs以英语为内部枢纽语言提供新证据。
AI 中文摘要
研究背景:已有研究表明,多语言大语言模型(LLMs)在非英语分类任务上的表现,与模型内给定语言的表示和英语的对齐程度正相关;目前已提出多种适用于LLMs的跨语言对齐(CLA)分数,以及多种从模型中提取词嵌入的方法。研究问题:现有工作几乎仅聚焦于CLA分数对分类任务的预测能力,而LLMs也广泛用于机器翻译等生成任务,因此需探究CLA分数是否同样能预测翻译性能。方法:为实现跨目标语言的相关性计算,提出一种基于PMI的翻译指标,该指标对目标语言的依赖程度更低,且与chrF指标具有强相关性;对27种CLA分数变体进行对比分析,考察其差异及对三类下游任务性能的预测效果。结果:与源-目标CLA相比,使用英语的CLA对翻译质量的预测能力相当或更优。结论:该结果为LLMs将英语用作内部枢纽语言提供了新的证据。
英文摘要
Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the given language are more aligned to English within the model. Several cross-lingual alignment (CLA) scores have been proposed for use with LLMs, along with multiple approaches for extracting embeddings from the models. We provide a comparative analysis of 27 CLA score variants, examining how they differ and how well each predicts downstream performance across three tasks. Crucially, while LLMs are widely used for generative tasks such as machine translation, prior work has focused almost exclusively on classification. We therefore investigate whether CLA scores are similarly predictive of translation performance. To enable computing correlations across target languages, we propose a PMI-based translation metric, which is less dependent on the target language and correlates strongly with chrF. We find that CLA with English predicts translation quality comparably to or better than source-target CLA, providing new evidence that LLMs use English as an internal pivot language.
CommentsAccepted to EMNLP 2026