arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32519cs.AIcs.LG

STR:用于减少转向副作用的监督式转码器替换

STR: Supervised Transcoder Replacement for Reducing Steering Side Effects

Haonan Yu, Junhao Liu, Zhenyu Yan, Haoran Lin, Xin Zhang

AI总结:

本文提出监督式转码器替换(STR)方法,通过监督学习替换转向层MLP计算,在保持目标控制的同时减少转向副作用,实验显示在Gemma模型上显著降低分布外攻击成功率。

AI中文摘要:

模型转向可以增强目标行为,但同时会削弱其他有用行为。我们引入了监督式转码器替换(STR)来减少现有转向方法(包括那些未配备保护目标的转向方法)的这些副作用。STR通过监督学习在转向层学习多层感知器(MLP)计算的替换,监督目标包括目标控制、非目标保留以及无转向时的保真度。选定的转向方法随后在冻结的替换上拟合方向,同时保留其自身的拟合目标。我们使用Corrigibility偏好和四个有害请求安全数据集,在Gemma和Llama模型上评估了三种转向方法。SALAD-Bench提供保护训练数据和独立的分布内评估分割;HarmBench、AdvBench和StrongREJECT保留用于分布外测试。STR显著减少了分布内评估中的转向副作用,并将这种保护扩展到未见过的安全数据集,同时保持有效的目标控制。对于仅目标监督转向向量,Gemma-3-4B上的汇总分布外攻击成功率从42.46%降至14.42%,Gemma-3-12B上从34.97%降至12.91%。这些结果表明,替换训练可以使未配备保护目标的转向方法受益。

英文摘要:

Model steering can strengthen a target behavior while degrading other useful behaviors. We introduce Supervised Transcoder Replacement (STR) to reduce these side effects for existing steering methods, including those fitted without a protection objective. STR learns a replacement for the multilayer perceptron (MLP) computation at the steering layer through supervision for target control, non-target preservation, and fidelity without steering. Selected steering methods then fit directions on the frozen replacement while retaining their own fitting objectives. We evaluate three steering methods across Gemma and Llama models using Corrigibility preferences and four harmful-request safety datasets. SALAD-Bench supplies protection training data and a separate in-distribution evaluation split; HarmBench, AdvBench, and StrongREJECT are reserved for out-of-distribution testing. STR substantially reduces steering side effects on the in-distribution evaluation and extends this protection to the unseen safety datasets while retaining effective target control. For target-only supervised steering vectors, pooled out-of-distribution attack success rate falls from 42.46% to 14.42% on Gemma-3-4B and from 34.97% to 12.91% on Gemma-3-12B. These results show that replacement training can benefit steering methods fitted without protection objectives.

↑