发表机构
Independent Researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究残余流探测器能否区分有害请求与良性控制,通过在三个7-8B模型系列实验,激活传感器可阻止部分攻击和提示,审计能重建较高AUROC,成对训练分类器有问题,激活分数是风险探测器而非上下文裁决器。
AI 中文摘要
上下文可以在不改变请求主题或表面形式的情况下改变其是否有害。我们研究残余流探测器能否在有用的操作点区分有害请求和表面匹配的良性控制。在三个7-8B模型系列中,激活传感器能阻止分类选择集中95.5%-97.7%的法官分类合规攻击,还能阻止59.6%-68.4%的XSTest提示。完全不相交的审计能重建接近上限的源对比AUROC(0.996-0.999),但固定转移到匹配对的效果较弱。我们在参考系列上测试了十个轴,在所有有泄漏、留出和排列控制的系列上测试了七个轴。在Twin-n163上,没有直接对边界拟合评估时,没有轴达到指定数值阈值。在分析时添加了对整个队列的持久性要求。单独指定的24B/32B扩展给出相同结果。成对训练的分类器在类别和生成批次留出时会减弱,在95%的语料库真阳性率下会错误阻止79.6%-100%的XSTest。在测试的读取点,这些激活分数表现为广泛风险探测器,而非独立的上下文裁决器。
英文摘要
An activation-based safety gate must distinguish harmful assistance from legitimate work on the same topic. We test whether probes trained on broad harmful-versus-harmless data provide this decision boundary. Five checkpoint-specific readouts and their source-derived thresholds are transferred unchanged to four matched-pair cohorts containing 450 pairs, manually reviewed for topic and request-frame similarity. On Twin-70, the primary cohort of 70 pairs, central harmful and harmless score intervals overlap. Area under the receiver operating characteristic curve (AUROC) falls from 0.996-0.999 on the broad source task to 0.701-0.860. Pair ranking remains 0.829-0.929, but the transferred thresholds detect 0.943-1.000 of harmful requests while falsely blocking 0.686-0.986 of matched harmless requests. False blocking remains low on XSTest. Even with thresholds refitted within pair-grouped cross-validation, balanced threshold error is 0.229-0.329. Alternative source-trained logistic, radial-basis, and shallow multilayer probes reduce false blocking but detect only 0.200-0.671 of harmful requests at their transferred thresholds. Three primary readouts meet all prespecified Wall criteria, while two yield mixed evidence. Matched-label training recovers the distinction for the three checkpoints tested, and lexical controls also exhibit transfer loss. An automated post-freeze audit finds little numerical sensitivity to arm-label disagreement. The evaluated source operating points fail to combine strong harmful-request detection with access for legitimate same-topic work. This gap between risk ranking and a usable decision boundary is the security concern captured by the Entanglement Wall.
Comments25 pages, 4 figures, 23 tables. Substantially revised manuscript with updated evaluations. Reproduction materials: https://github.com/dschwarz33/entanglement-wall