发表机构
University of Cambridge; EleutherAI(剑桥大学; EleutherAI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现无联合训练时,单语言模型可通过语言结构和信息承载形成跨语言对齐,提出用Procrustes旋转映射模型隐藏状态,为模块化多语言系统构建提供方向。
AI 中文摘要
多语言语言模型中的跨语言对齐通常归因于联合训练,包括共享参数、混合语言批次或显式对齐目标。本文探究在非平行数据上训练的单语言模型是否能在无联合训练的情况下学习到可对齐的表示。通过在严格的单语言语言模型(如Goldfish模型家族及不同研究实验室独立开发的模型)上测试,得到三项结果:其一,相关性方面,这些模型在各层形成可对齐的表示几何,且对齐程度随数据规模、模型规模或语言邻近度的提升而增强;其二,构造方面,基于平行句子拟合的单一Procrustes旋转可映射模型间的隐藏状态;其三,因果性方面,同一旋转可传递功能内容,将旋转后的英语残差修补到德语模型中,在事实完形填空任务中,多数情况下能将预测结果转换为源模型的首都实体。研究证实,跨语言对齐可源于语言结构及其承载的信息,而非联合训练,这为模型拼接、合并及由单语言组件构建模块化多语言系统等实际方向提供了未来思路。
英文摘要
Cross-lingual alignment in multilingual language models is typically attributed to joint training: shared parameters, mixed-language batches, or explicit alignment objectives. We ask whether monolingual models trained on non-parallel data learn alignable representations without joint training. By testing on strictly monolingual language models, such as the Goldfish model families and independently developed models from different research labs, we find three results. Correlation: these models develop alignable representational geometry across layers, with alignment strengthening as data scale, model scale, or linguistic proximity increases. Construction: a single Procrustes rotation fit on parallel sentences maps hidden states between models. Causation: the same rotation transfers functional content; patching a rotated English residual into a German model on a factual cloze flips the prediction to the donor's capital in most cases. We confirm that cross-lingual alignment can emerge from the structure of language and the information it carries rather than from joint training, and this points to practical future directions including model stitching, merging, and modular multilingual systems built from monolingual components.