arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一路相似:大语言模型中的多语言泛化依赖于语言层面的相似性结构

Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures

Supantho Rakshit, Adele Goldberg, Henry Conklin

arXiv 2607.22699首次发表:更新:

发表机构

Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型多语言泛化能力,通过借鉴认知科学,探究其对不同语言层次相似性结构的捕捉,发现潜在表示能恢复印欧语系语言家族树结构,且模型反映语言相似性程度与XNLI表现相关,扩展了相似性驱动泛化工作。

AI 中文摘要

随着大语言模型(LLMs)在各种任务中能力不断增强,其泛化能力难以量化且在有限领域之外理解不足。特别是在多语言泛化方面,对英语以外且在训练数据中出现较少的语言存在困难。为探究原因及模型表现差异,我们借鉴认知科学的研究,认为成功泛化源于相似性空间中的适当表示。我们研究LLMs的表示如何捕捉不同语言间的层次相似性结构。惊人的是,我们发现LLMs的潜在表示很大程度上恢复了印欧语系语言家族树的层次结构,且模型反映语言相似性结构的程度与它们在多语言自然语言推理基准XNLI上的表现相关。这扩展了大规模相似性驱动泛化的经典工作,表明相似语言表示相似的模型在跨语言泛化上表现更好。

英文摘要

As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and poorly understood beyond limited domains. In particular, LLMs are known to struggle generalizing multilingually, to languages outside of English, and that are poorly attested in their training data. To understand why this may be, and what enables some models to perform better than others, we turn to a long history of work across the cognitive sciences, arguing that successful generalization derives from appropriate representations in similarity space. We look at how well LLMs' representations capture the hierarchical similarity structure between distinct languages. Strikingly, we show LLMs' latent representations largely recover the hierarchical structure of the Indo-European language family tree -- grouping languages that are members of the same subfamily closely together in representation space. Furthermore, we show that the degree to which models reflect the similarity structure of languages correlates with their performance on XNLI, a multilingual natural language inference benchmark. This extends classic work on similarity-driven generalization at scale, showing how models that represent similar languages similarly generalize better from one language to another.

CommentsAccepted for oral presentation at CogSci 2026 (48th Annual Meeting of the Cognitive Science Society), Rio de Janeiro. 8 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑