发表机构
Gödel Machines; IIT Madras(哥德尔机器; 印度马德拉斯理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多语言MoE模型的跨语言拒绝不均衡问题,在sarvam模型中定位到由注意力对抗者约束的混合专家生成器回路,揭示其机制并量化了安全修复的成本。
AI 中文摘要
多语言模型的安全对齐存在不均衡问题:在英语中能可靠拒绝有害请求的模型,在低资源语言中往往会顺从相同请求。我们在印度多语言混合专家推理模型sarvam中从机制层面追溯这一差距,发现这并非是无法检测到危害。危害被编码为网络中层(第11层)几乎与语言无关的内部方向(英语与印度语言的余弦相似度≈0.9),上游调控该方向可因果控制拒绝行为。但检测方向与实际生成拒绝的变化正交,该变化是滞后的,在生成过程中逐步组装,而非通过单次前向传播读取。我们将这种拒绝生成归因于特定、可定位的回路:由注意力对抗者约束的混合专家生成器。我们测试了干预该回路的所有方式:抑制对抗者成本低且有效,放大生成器存在成本壁垒,对负责的注意力头进行精准编辑则无效。该回路的结构及揭示它的梯度方法在第二个无关的MoE模型中也存在,而调控杠杆的强度具有架构特异性。研究最终得到了多语言安全修复可落地的位置及对应成本的量化图谱。
英文摘要
Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in a lower-resource language. We trace this gap mechanistically in sarvam, an Indic-multilingual mixture-of-experts reasoning model, and find it is not a failure to detect harm. Harm is encoded as an internal direction that is nearly language-invariant in mid-network (English-vs-Indic cosine ${\approx}0.9$ at $L11$), and steering that direction upstream causally controls refusal. But the detection direction is orthogonal to the change that actually writes the refusal, which is late and assembled over the course of generation rather than read off in a single forward pass. We attribute the write to a specific, localizable circuit, a mixture-of-experts writer held in check by an attention opposer and price every way of intervening on it: damping the opposer is cheap and effective, amplifying the writer is a cost wall, and surgical edits to the responsible heads do nothing. The circuit's organization, and the gradient method that exposes it, recur in a second, unrelated MoE model, while the lever's strength is architecture-specific. The result is a cost-measured map of where a multilingual safety repair can land, and what it costs
CommentsAccepted to the actionable Interpretability workshop at COLM 2026