多语言大语言模型中从表征到输出的刻板印象追踪
Tracing Stereotypes from Representation to Output in Multilingual LLMs
浏览论文内容
中文总结 AI 辅助
本文通过探针、归因修补和稀疏自编码器等方法,追踪多语言大模型中刻板印象的表征位置与输出影响,发现探针早于归因且跨语言效应有限。
中文摘要 AI 辅助
多语言大语言模型表现出因语言而异的刻板印象相关行为,但行为分数无法显示相关信息在何处被表征或如何影响输出。为探究这些内部机制,我们在Llama-3.1-8B、Qwen3-8B和Gemma-2-9B中比较了线性探针、归因修补、稀疏自编码器(SAEs)和特征消融。在所有三个模型中,探针性能的峰值显著早于归因,两者间隔占模型深度的36%-53%。保留的Llama-Scope特征通常与其被选择时所依据的社会类别相匹配,并形成重复出现的语义族,但其词汇对齐和消融效果在不同SAE套件间存在差异。在我们的标准下,仅6%-18%的评估残差流特征具有语言无关效应,且没有特征具有类别无关效应。语言无关特征在Llama-Scope中具有更大的平均消融效应,但这一模式在其他SAE套件中并未重复。因此,可解码性、输出影响和跨语言消融效应需要分别测量。
英文摘要
Multilingual LLMs show stereotype-related behavior that varies across languages, but behavioral scores do not show where the relevant information is represented or how it affects the output. To investigate these internal mechanisms, we compare linear probing, attribution patching, sparse autoencoders (SAEs) and feature ablation in Llama-3.1-8B, Qwen3-8B, and Gemma-2-9B. Probe performance peaks substantially earlier than attribution in all three models, with a separation of 36-53% of model depth. Retained Llama-Scope features often match the social category on which they were selected and form recurring semantic families, but their lexical alignment and ablation effects vary across SAE suites. Only 6-18% of evaluated residual-stream features have language-agnostic effects under our criterion, and none are category-agnostic. Language-agnostic features have larger mean ablation effects in Llama-Scope, but this pattern does not repeat in the other SAE suites. Decodability, output influence, and cross-lingual ablation effects therefore need to be measured separately.
发表机构
- Saarland University(萨尔大学)
- Queen’s University Belfast(贝尔法斯特女王大学)
- German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI))
- University of Technology Nuremberg(纽伦堡工业大学)
机构由 AI 辅助整理,请以论文原文为准。