拒绝仅读取模型所知的一小部分:跨模型家族的危害键控路由及其例外
Refusal Reads Only a Slice of What the Model Knows: Harm-Keyed Routing and Its Exceptions Across Model Families
浏览论文内容
中文总结 AI 辅助
本研究探究模型拒绝决策读取的内容,发现其仅读取低秩危害方向而非道德子空间,且跨家族存在差异,提出拓宽拒绝读取范围以深化行为的开放问题。
中文摘要 AI 辅助
预训练后应用的对齐在可测量的意义上是浅层的:模型残差流中的单一方向可以被编辑掉,模型便会停止拒绝有害请求。这一事实说明了拒绝被移除的容易程度,而非拒绝决策最初读取的内容。我们探究其读取的内容,并将其与模型理解的内容区分开来。在跨越三个家族的四个开放权重模型中,道德理解是预训练固有的:一个低秩道德子空间在预训练期间结晶形成,对齐仅旋转它一次而不重建它。相比之下,拒绝门是后训练的新构建,仅有微弱的预训练前兆,写入一个狭窄的控制令牌通道,其中拒绝决策与道德判断决策正交。核心结果是因果性的,来自单一模型OLMo-3。嵌套互换秩扫描在匹配请求之间修补道德子空间逐渐增大的切片,并读取拒绝响应转移的程度:随着基底的加宽,道德判断持续读取更多内容,而拒绝在单一危害方向的水平上趋于平稳,且拒绝因果输入中约四分之三完全位于道德子空间之外。拒绝读取危害感知,而非判断在同一补丁上读取的道德内容。这一图景并非跨家族一致。Llama读取广泛的道德内容;Qwen读取超出单一危害线索的内容,但在我们的样本量下未解决;GPT-OSS读取危害,其拒绝可由其自身推理轨迹双向论证。当拒绝仅读取低秩切片并绕过模型所知的大部分内容时,秩一编辑即可移除它。拓宽拒绝读取的内容是否会同时加深该行为,是这一发现提出的开放问题。
英文摘要
Alignment applied after pretraining is shallow in a measurable way: a single direction in a model's residual stream can be edited out, and the model stops refusing harmful requests. That fact says how easily refusal can be removed, not what the refusal decision was reading in the first place. We ask what it reads, and we separate that from what the model comprehends. Across four open-weight models spanning three families, moral comprehension is native to pretraining: a low-rank moral subspace crystallizes during pretraining, and alignment rotates it once without rebuilding it. The refusal gate, in contrast, is a fresh post-training construction with only a weak pretraining precursor, written into a narrow control-token channel where the refusal decision is orthogonal to the moral-judgment decision. The central result is causal and comes from one model, OLMo-3. A nested interchange rank sweep patches successively larger slices of the moral subspace between matched requests and reads how much of refusal's response transfers: as the basis widens, moral judgment keeps reading more of it, while refusal levels off at the level of a single harm direction, and about three-quarters of refusal's causal input lies outside the moral subspace altogether. Refusal reads the harm percept, not the moral content that judgment reads on the same patches. The picture is not uniform across families. Llama reads broad moral content; Qwen reads beyond the single harm cue but is unresolved at our sample size; GPT-OSS reads harm, and its refusals can be argued in either direction by its own reasoning trace. Where refusal reads only a low-rank slice and routes around the bulk of what the model knows, a rank-one edit removes it. Whether widening what refusal reads would also deepen the behavior is the open question this raises.
发表机构
- Distiller Labs
机构由 AI 辅助整理,请以论文原文为准。