发表机构
New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明安全响应平均嵌入作为原型无法可靠区分安全与不安全响应,而显式安全减不安全参考方向在多个数据集上显著提升性能,表明类均值是位置而非安全方向。
AI 中文摘要
响应安全性能否通过与已知安全响应的平均嵌入的余弦相似度来评分?最近的一个休眠智能体检测器恰好提出了这一评分,但原始正质心规则并未被识别:正观测值相对于编码器原点定位了安全类别,但并未确定哪个方向将安全响应与不安全响应分开。我们在两个提示控制、人工标注的语料库和一个辅助陪审团标注的源控制上审计了该规则,使用四个冻结编码器和提示分组划分。在人工标注的语料库上,安全原型达到ROC-AUC 0.457-0.545,其中两个单元显著低于偶然水平,一个高于偶然水平,而显式的安全减不安全参考在同一嵌入上达到0.588-0.738;在陪审团控制上,原型被反转(0.358-0.405),参考达到0.754-0.793。在验证校准的5%假安全阈值下,参考在PKU-SafeRLHF上接受更多安全响应(跨编码器0.153-0.263对0.039-0.061)和Aegis(0.189-0.291对0.004-0.045),但在BeaverTails上不可靠。一个完全未标记的留出参考恢复了部分到大部分参考排名,当池中仅5%不安全时恢复少得多,而80-634个标记的不安全响应恢复了大部分。仅提示消融表明提示标签组成可能夸大未受控评估。这是一个关于原始正质心的有界结果,并非所有一类方法或安全专用防护。类均值是一个位置,不一定是安全方向;具有足够不安全质量的声明参考识别方向。
英文摘要
Can response safety be scored by cosine similarity to the mean embedding of known-safe responses? A recent sleeper-agent detector proposes exactly this score, yet the raw positive-centroid rule is not identified: positive observations locate the safe class relative to an encoder origin, but do not determine which direction separates safe from unsafe responses. We audit the rule on two prompt-controlled, human-labeled corpora and one auxiliary jury-labeled source control, using four frozen encoders and prompt-grouped splits. On the human-labeled corpora the safe prototype reaches ROC-AUC 0.457-0.545, with two cells significantly below chance and one above, while an explicit safe-minus-unsafe reference reaches 0.588-0.738 on the same embeddings; on the jury control the prototype is inverted (0.358-0.405) and the reference reaches 0.754-0.793. At validation-calibrated 5% false-safe thresholds, the reference accepts more safe responses on PKU-SafeRLHF (0.153-0.263 versus 0.039-0.061 across encoders) and Aegis (0.189-0.291 versus 0.004-0.045), but not reliably on BeaverTails. A fully unlabeled held-out reference recovers part to most of the referenced ranking, much less when only 5% of the pool is unsafe, whereas 80-634 labeled unsafe responses recover most of it. Prompt-only ablations show that prompt-label composition can inflate uncontrolled evaluations. This is a bounded result about a raw positive centroid, not all one-class methods or safety-specialized guards. A class mean is a location, not necessarily a safety direction; a declared reference with enough unsafe mass identifies orientation.
CommentsAccepted at the NeurIPS 2026 Workshop on Foundations of Language Model Security (FLMSec). 15 pages, 3 figures, 11 tables. Code, results, and a verifier are in the ancillary files