大语言模型中的文字选择:晚期层承诺的证据
Script Choice in LLMs: Evidence for Late-Layer Commitment
浏览论文内容
中文总结 AI 辅助
本文通过逻辑回归探针和logit-lens分析,发现大语言模型对输出文字的承诺发生在最后层,且与模型深度相关,为多语言架构设计提供启示。
中文摘要 AI 辅助
本文利用两种互补的可解释性方法——逻辑回归探针和logit-lens分析——研究文字知识如何分布在大语言模型(LLM)的各层中。我们的探针实验揭示了一个明显的不对称性:输入文字和指示的输出文字都在网络的最早层中被编码,而相比之下,对实际输出文字的承诺仅出现在最后几层,模型在大部分层的中间表示默认使用拉丁文字。这一两阶段过程得到了logit-lens分析的证实,该分析显示文字承诺始终发生在LLM的最后几层。结合在较小模型中观察到的较弱的文字遵循性能,这些结果形成了一致的证据链,将文字承诺与模型深度联系起来,并对设计足够深、包容性的多语言架构具有更广泛的意义。
英文摘要
In this paper, we investigate how script knowledge is distributed across the layers of LLMs using two complementary interpretability methods: logistic regression probing and logit-lens analysis. Our probing experiments reveal a clear asymmetry: both the input script and the instructed output script are encoded in the earliest layers of the network, while, in contrast, commitment to the actual output script emerges only in the final layers, with the model's intermediate representations defaulting to Latin throughout most of the layers. This two-stage process is confirmed by logit-lens analyses, which show that script commitment consistently occurs at the very last layers of the LLMs. Together with the weaker script-following performance observed in smaller models, these results form a converging body of evidence linking script commitment to model depth, with broader implications for the design of sufficiently deep, inclusive multilingual architectures.
发表机构
- SUPSI, IDSIA, Switzerland(瑞士南部应用科学与艺术大学,IDSIA)
- armasuisse, Science & Technology, Switzerland(瑞士联邦国防采购局,科学与技术部)
机构由 AI 辅助整理,请以论文原文为准。