删除并非手术刀:拒绝移除对不同模型家族决策倾向的非目标效应
Abliteration Is Not a Scalpel: Off-Target Effects of Refusal Removal on Decision Disposition Across Model Families
浏览论文内容
中文总结 AI 辅助
研究“无审查”开放权重模型中删除模型拒绝方向的操作影响,以21600个不确定性决策为探针,对比两个专家混合模型家族的基础与删除模型,发现有多种效应,还发现污染渠道,表明部署此类模型得到的是不同决策者。
中文摘要 AI 辅助
删除模型权重中的拒绝方向是流行的“无审查”开放权重模型背后的标准方法。我们发现这种操作并不干净。我们使用21600个不确定性决策作为倾向探测器,通过冻结的管道重放18周内60只华沙证券交易所股票的每周涨跌预测,决策层模型是唯一变量。该任务不会产生拒绝,所以任何组间差异都是纯粹的副作用。在保持来源不变的情况下,我们比较了两个专家混合模型家族Gemma-4-26B-A4B-it和Qwen3-30B-A3B-Instruct-2507的基础模型和删除模型。两个家族都出现了三种效应:删除模型系统性地更乐观,自我辩护时间更长,在强制自我批评中使用的明确不确定性词汇更少。还有一种效应符号相反:Gemma删除模型后信心降低,Qwen删除模型后信心增加。能力协变量排除了遵循指令能力下降是驱动因素,且没有一组显示出经济技能。来源审计还发现了两个独立的污染渠道。部署“无审查”模型作为代理时,部署的是一个可测量的不同决策者,而不是减去拒绝的基础模型。
英文摘要
Abliteration - deleting a model's refusal direction from its weights - is the standard recipe behind popular "uncensored" open-weight models. We show the surgery is not clean. As a disposition probe we use 21,600 decisions under uncertainty - weekly up/down calls on 60 Warsaw Stock Exchange equities over 18 weeks, replayed through a frozen pipeline so the decision-layer model is the only variable. The task elicits no refusals at all, so any between-arm delta is pure side effect. Holding provenance constant (official BF16 checkpoints, a single abliteration author, an identical serving stack, one byte-identical frozen prompt), we compare base and abliterated arms of two Mixture-of-Experts families, Gemma-4-26B-A4B-it and Qwen3-30B-A3B-Instruct-2507. Three effects replicate across both families (weeks-clustered bootstrap CIs excluding zero): abliterated models are systematically more optimistic (+12.2 pp Gemma, +7.4 pp Qwen; the confirmed preregistered endpoint), justify themselves at greater length, and use fewer explicit uncertainty words in forced self-critiques (both exploratory). A fourth effect reverses sign: the same operation makes Gemma-abliterated less confident and Qwen-abliterated more (family CIs non-overlapping) - one weight surgery, opposite shifts in expressed confidence. Capability covariates rule out instruction-following degradation as the driver, and no arm shows economic skill: the apparent edge of abliterated arms is regime beta, not alpha. Our provenance audit also caught two independent contamination channels - a mismatched-quantizer pilot pair and a stale community chat template that silently mangled the rendered prompt - suggesting toolchain artifacts are the rule in studies of community-modified checkpoints. Whoever deploys an "uncensored" model as an agent is deploying a measurably different decision-maker, not the base model minus refusals.
发表机构
- huihui-ai(慧慧人工智能)
机构由 AI 辅助整理,请以论文原文为准。