利用几何不变稀疏自编码器发现大语言模型中的跨语言推理不变性
Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders
浏览论文内容
中文总结 AI 辅助
本研究以MGSM数据集探究多语言LLM跨语言推理的特征机制,提出GI-SAE方法,发现跨语言特征共享依赖模型架构,GI-SAE主要放大现有跨语言结构且模型特异性明显。
中文摘要 AI 辅助
多语言大语言模型能够用不同语言求解同一数学问题,但目前尚不清楚它们是依赖共享特征,还是依赖仅产生相似输出的特定语言计算。本研究使用Multilingual Grade School Math(MGSM)数据集,在四个系列的五个模型中探究该问题,数据集包含用英语、德语、法语、西班牙语、俄语和中文求解的问题,保留所有六种语言中具有有效推理轨迹的问题,并通过模型重放这些轨迹以记录多层的表示。对于每个模型,首先使用中心化核对齐(Centered Kernel Alignment, CKA)识别具有跨语言对齐的层。在每个选定层,训练两个稀疏自编码器(Sparse Autoencoders, SAE):一个仅用于重建的基线模型,以及本研究提出的对比变体——几何不变稀疏自编码器(Geometry-Invariant SAE, GI-SAE)。GI-SAE在重建损失基础上补充了信息噪声对比估计(Information Noise-Contrastive Estimation, InfoNCE)损失,该损失训练编码器为同一问题的推理轨迹生成相似激活,无论语言或标记位置如何。随后通过在模型前向传播过程中交换不同语言间的特征值,并测量输出变化(以每个特征的Kullback-Leibler, KL散度量化),测试所得共享特征是否可功能互换。尽管GI-SAE在几乎所有层都产生了更高的CKA和Jaccard相似度,但更高的几何相似度并不始终意味着更强的功能互换性。研究发现,本样本中的跨语言特征共享依赖于模型和架构,且在不同模型中出现在不同深度;GI-SAE主要放大已存在的跨语言结构,该模式具有模型特异性:在Qwen模型中得到增强,在Gemma模型中无功能收益,在Llama和Phi模型中则存在与层相关的混合效应。
英文摘要
Multilingual language models can solve the same mathematical problem in different languages, but it remains unclear whether they rely on shared features or on language-specific computations that only produce similar outputs. We study this question in five models from four families using the Multilingual Grade School Math (MGSM) dataset, with problems solved in English, German, French, Spanish, Russian, and Chinese, retaining problems with valid reasoning traces in all six languages and replaying those traces through the model to record representations at multiple layers. For each model, we first use Centered Kernel Alignment (CKA) to identify layers with cross-language alignment. At each selected layer, we train two sparse autoencoders (SAE): a baseline reconstruction-only model and a contrastive variant introduced in this work, the Geometry-Invariant SAE (GI-SAE). GI-SAE supplements the reconstruction loss with an Information Noise-Contrastive Estimation (InfoNCE) loss that trains the encoder to produce similar activations for traces of the same problem, regardless of language or token position. We then test whether the resulting shared features are functionally interchangeable by swapping their values between languages during the model's forward pass and measuring the resulting change in output, quantified by Kullback-Leibler (KL) divergence per feature. Although GI-SAE yields higher CKA and Jaccard similarity at nearly every layer, higher geometric similarity does not consistently imply greater functional interchangeability. We find that cross-language feature sharing is model- and architecture-dependent in this sample and appears at different depths in different models. GI-SAE primarily amplifies cross-language structure already present: the pattern is model-specific, with strengthening in Qwen, no functional benefit in Gemma, and mixed layer-dependent effects in Llama and Phi.
发表机构
- Carleton University(卡尔顿大学)
机构由 AI 辅助整理,请以论文原文为准。