发表机构
EPFL(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究多语言大语言模型中不同潜在语言探测的差异,发现两类探测结果分歧,提示需谨慎解读潜在语言识别,其揭示多语言处理的不同方面而非单一内部通用语。
AI 中文摘要
潜在语言识别常被用于论证多语言语言模型通过特定语言状态(如英语枢纽)传递计算。然而,现有探测方法从不同信号推断潜在语言,如隐藏状态的几何结构或可从中间表示解码的内容。由于此类结论关乎模型如何跨语言共享与传递信息,我们探究这些探测是否测量同一现象,或揭示多语言计算的不同方面。我们在不同模型家族、训练范式、领域、任务、检查点及多达27种语言上研究该问题。发现识别探测存在系统性分歧:基于隐藏状态几何结构的GMM表示探测显示更早的跨语言混合,而依赖输出空间可解码性的解码探测保留更清晰的语言特异性及更偏向英语的信号。这些差异与模型多语言性及训练进程相关,但跨领域相对稳定。结果提示需更谨慎解读潜在语言识别,当前探测揭示多语言处理的不同方面,而非直接揭示单一内部通用语。
英文摘要
Latent language identification is often used to argue that multilingual language models route computation through language-specific states, such as English pivots. However, existing probes infer latent language from different signals, such as the geometry of hidden states or what can be decoded from intermediate representations. Since such claims shape conclusions about how models share and route information across languages, we ask whether these probes measure the same phenomenon or expose distinct aspects of multilingual computation. We study this question across model families, training regimes, domains, tasks, checkpoints, and up to 27 languages. We find that identification probes systematically disagree: the GMM-based representation probe, which draws evidence from hidden state geometry, shows earlier cross-lingual mixing, whereas decoding-based probes, which rely on output-space decodability, retain sharper language-specific and more English-biased signals. These differences track model multilinguality and training progression, but are comparatively stable across domains. Our results suggest a more cautious interpretation of latent language identification, where current probes expose different aspects of multilingual processing, rather than directly revealing a single internal lingua franca.
CommentsAccepted to EMNLP 2026 Findings