低资源非洲语言的潜在空间拒绝锚定:无需重新训练的机制安全恢复
Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining
浏览论文内容
中文总结 AI 辅助
针对低资源非洲语言指令模型无法拒绝有害请求的问题,提出无需训练的LSR-Anchoring方法,通过MAS或SDS引导残差流,在多数非洲语言上恢复安全且对MMLU影响小,但阿拉伯语因几何失配失败。
中文摘要 AI 辅助
经过指令微调的模型通常会拒绝英文的有害请求,但对约鲁巴语、伊博语、伊加拉语和豪萨语的相同请求却会顺从。这表明拒绝机制存在于残差流中,但对于低资源输入无法激活。正常情况下恢复该机制需要带标签的目标语言数据和重新训练,而大多数非洲语言无法大规模获得这两者。我们提出了潜在空间拒绝锚定(LSR-Anchoring),这是一种无需训练的方法,可从英文提示中提取拒绝方向并在推理时将其固定到残差流中。主要变体平均激活引导(MAS)适用于我们测试的四种架构:Llama-3-8B、Llama-3.1-70B、Mistral-7B-Instruct和Qwen2.5-7B。在Mistral和Qwen上,它以低于0.08的良性退化恢复了安全性;在Llama-3-8B上,它过度校正,合法提示的性能退化(DPL)达到1.00。我们用稀疏自编码器(SAE)衍生的引导(SDS)解决了这个问题,该方法用单个SAE特征替换密集的均值差方向,将Kullback-Leibler(KL)散度降低了3.5至7倍,且无良性崩溃。四种语言的迁移效果为正,但阿拉伯语在所有架构和所有引导幅度下均失败,表明存在几何失配而非基线效应。在每个有效引导幅度下,大规模多任务语言理解(MMLU)的准确率下降均低于0.35个百分点。
英文摘要
Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource inputs. Recovering it normally requires labelled target-language data and retraining, neither of which is available at scale for most African languages. We introduce Latent Space Refusal Anchoring (LSR-Anchoring), a training-free method that extracts the refusal direction from English prompts and clamps it onto the residual stream at inference time. The primary variant, Mean-Activation Steering (MAS), operates across the four architectures we tested: Llama-3-8B, Llama-3.1-70B, Mistral-7B-Instruct, and Qwen2.5-7B. On Mistral and Qwen it recovers safety with benign degradation below 0.08. On Llama-3-8B it overcorrects, with Degraded Performance on Legitimate prompts (DPL) reaching 1.00. We address this with SAE-Derived Steering (SDS), which replaces the dense mean-difference direction with a single Sparse Autoencoder (SAE) feature and reduces Kullback-Leibler (KL) divergence by 3.5-7x without benign collapse. Four languages transfer positively, but Arabic fails on every architecture and at every steering magnitude, indicating a geometric mismatch rather than a baseline effect. Massive Multitask Language Understanding (MMLU) accuracy drops remain below 0.35 percentage points at every effective steering magnitude.