arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04316cs.AIcs.CL

从潜空间到雅可比空间:测量、规避与训练对抗安全内容可访问性

From Latent Space to Jacobian Space: Measuring, Evading, and Training Against Safety-Content Accessibility

Mohammad Mosafer

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过雅可比空间测量模型内部安全内容的可访问性,提出传输-放大剖面与配对协议,揭示安全训练可扩大可访问性差距,并强调监控、训练与评估应针对可访问性本身。

中文摘要 AI 辅助

仅输出端的安全监控只能看到模型计算的末端,但模型在输出之前就已计算出其答案:它在内部准备说的内容具有安全关键性。雅可比空间(J-space)读出,即通过模型的输入-输出雅可比矩阵从隐藏状态到输出词汇表的线性映射,已被提出作为行为上可访问的内部内容的窗口,并且首个安全协议(JADR)表明危险识别在该空间是可读的。目前仍未知的是,这种可访问内容与表示工程所研究的潜空间安全几何有何关联,以及安全训练如何塑造它。我们引入了两个定量桥梁:(i)传输-放大剖面 $A_\ell$,衡量每个层的雅可比矩阵将潜在安全方向携带到输出空间的强度,以及(ii)配对的基础-微调协议,将可访问性归因于训练来源。在四个小型模型对(135M至1.5B,三个透镜拟合种子)中,微调检查点将识别集中在中间偏上层,而在0.5B规模下,DPO安装拒绝行为的同时,J空间识别降至随机水平:训练可以在行为安全看起来最佳之处恰好扩大可访问性差距。部署审计闭环:监控器感知的GCG式后缀在每个规模下以零行为成本抑制提示侧监控器;只有学习到的监控器层级能抵抗自身的自适应再攻击;训练时防御在每个惩罚权重下均在其再攻击中失败;引导显示放大剖面是描述性的而非因果性的;连续前缀优化在离散搜索成功之处失败。安全监控、训练和评估必须作用于可访问性本身,而非仅作用于输出。

英文摘要

Output-only safety monitoring sees only the end of a model's computation, yet the model computes its answer before emitting it: what it is internally poised to say is safety-critical. Jacobian-space (J-space) readouts, linear maps from hidden states to the output vocabulary via the model's input-output Jacobian, have been proposed as a window on behaviorally accessible internal content, and a first safety protocol (JADR) showed that danger recognition is readable there. What remains unknown is how this accessible content relates to the latent-space safety geometry studied by representation engineering, and how safety training shapes it. We introduce two quantitative bridges: (i) a transport-amplification profile $A_\ell$, measuring how strongly a latent safety direction is carried toward output space by each layer's Jacobian, and (ii) a paired base-vs-tuned protocol that attributes accessibility to training provenance. Across four small model pairs (135M to 1.5B, three lens-fit seeds), tuned checkpoints concentrate recognition in upper-middle layers, and at 0.5B DPO installs refusal while J-space recognition drops to chance: training can widen the accessibility gap exactly where behavioral safety looks best. A deployment audit closes the loop: monitor-aware GCG-style suffixes suppress the prompt-side monitor at zero behavioral cost at every scale; only a learned monitor rung resists its own adaptive re-attack; the training-time defense fails its re-attack at every penalty weight; steering shows the amplification profile is descriptive, not causal; and continuous-prefix optimization fails where discrete search succeeds. Safety monitoring, training, and evaluation must operate on accessibility itself, not on outputs alone.

补充信息

↑