操控语言轴:从线性可解码到因果控制
Steering the Language Axis: From Linear Decodability to Causal Control
浏览论文内容
中文总结 AI 辅助
本研究探究大语言模型的语言身份是否可通过紧凑激活方向因果控制,在Qwen、Llama等模型的126万次生成上证实,分离的PCA导出语言轴可可靠操控语言切换,且语言选择具层特异性与语言对依赖性。
中文摘要 AI 辅助
尽管大语言模型具备令人印象深刻的多语言能力,但决定语言选择的潜在动态机制仍鲜为人知。本研究探究语言身份是否仅能从隐藏状态中线性解码,抑或可通过紧凑的激活方向进行因果控制。我们在多个模型家族(包括Qwen 3.5-2B和Llama-3.2-1B-Instruct)上开展详尽的因果干预分析,分离出由PCA导出的“语言轴”,在FLORES-200数据集的126万次生成结果上进行操控与消融实验。沿这些几何方向操控可可靠地在跨脚本(英语到中文)和同脚本(英语到西班牙语)设置中强制语言切换,而同等幅度的随机扰动几乎无效果。我们的分层分析显示,语言选择具有高度局部性且明确依赖语言对:英语到中文的切换难以在早期层干预,在后期层则易于操控;而英语-西班牙语的切换更早发生,呈现出独特的双峰敏感性。此外,针对性消融揭示了一个基本的回归现象:一旦移除语言信号,无论输入提示如何,模型都会回归到英语。最终,这些发现表明,语言决策边界在推理过程中作为依赖方向且具有层特异性的因果活跃特征发挥作用。
英文摘要
Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly understood. In this work, we ask whether language identity is merely linearly decodable from hidden states, or if it can be causally controlled by a compact activation direction. We conduct an exhaustive causal intervention analysis across multiple model families, including Qwen 3.5-2B and Llama-3.2-1B-Instruct, isolating PCA-derived "language axes" to perform steering and ablation experiments across 1.26 million generations on the FLORES-200 dataset. Steering along these geometric directions reliably forces language switching in both cross-script (English to Chinese) and same-script (English to Spanish) settings, whereas equal-magnitude random perturbations yield virtually no effect. Our layerwise analysis reveals that language commitment is highly localized and explicitly language-pair-dependent. While English to Chinese switching resists early intervention and steers easily in the later layers, the English-Spanish transition shifts earlier, displaying a distinct, bimodal sensitivity. Furthermore, targeted ablation uncovers a fundamental reversion to English: once the language signal is removed, the model falls back to English regardless of the input prompt. Ultimately, these findings demonstrate that language decision boundaries function during inference as causally active features that are direction-dependent and layer-specific.
发表机构
- University of California, Santa Cruz(加利福尼亚大学圣克鲁兹分校)
机构由 AI 辅助整理,请以论文原文为准。