arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

召回并非保护:评估安全监控器对模型遵从性的效果

Safety Monitors Mostly Catch What the Model Already Refuses

Sripad Karne

arXiv 2609.05797首次发表:更新:

发表机构

Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现安全监控器在模型会遵从的有害提示上召回率显著更低,标准召回率高估了实际保护,应基于模型实际响应评估监控器。

AI 中文摘要

安全监控器对部署的语言模型接收的提示进行筛查,标记有害请求以确保其不被回答。这些监控器通常通过针对有害性标签的召回率来评估,但只有当模型原本会遵从请求时,拦截才能防止危害。我们直接测量这一差异:我们从目标模型中采样重复响应,如果模型至少一次遵从了有害提示,则将该提示称为“可诱发”,并分别报告监控器在可诱发和不可诱发提示上的召回率。在六种监控配置和三个模型家族中,涵盖激活探针、微调文本防护以及一个120B参数策略条件推理分类器,在固定假阳性率下,可诱发提示上的召回率比不可诱发提示上的召回率低0.22至0.38。监控器漏掉的提示被遵从的可能性是监控器捕获提示的2.8至5.6倍。这一差距在三个模型家族中重复出现,并且也出现在完全独立于目标模型的纯文本监控器中。这表明标准召回率可能高估了监控器在实际中提供的保护,监控器应针对其模型实际会回答的内容进行评估。

英文摘要

Safety monitors are evaluated by recall on harmful prompts, regardless of whether the target model would answer them. Yet a monitor matters most on the prompts the model does answer. We measure recall on exactly those prompts, defined by sampling the target model and judging its responses. Across four text guards, two activation probes, and Latent Guard, recall at a 1% false positive rate falls sharply on this subset: at a common threshold, every monitor catches the requests the model refuses 1.1 to 6.4 times as often as the requests it answers. Standard metrics hide this; AUROC stays above 0.85 for most monitors. Rewriting each request to be less explicit, with intent held fixed and verified, raises compliance 28-fold and lowers every monitor's flag rate. Of the requests newly answered after rewriting, 44 to 93% slip past the monitor, depending on which is used, and most of their completions are graded harmful. We trace the gap to explicitness itself. As wording softens with intent fixed, the target model's harm and refusal readings fall and it answers; the guards' harm readings fall too, and steering a guard along explicitness alone flips its verdict. Model and monitors miss the same prompts, and stacking monitors does not recover them. Fine-tuning a guard on the rewrites at every level of explicitness, on both sides of the label, raises recall on answered requests from .24 to .89 while transferring to unseen benchmarks; hard negatives, the natural alternative, teach the guard to discount indirect phrasing instead.

Commentsv3: substantially extended and retitled; v1 appeared as "Recall Is Not Protection: Evaluating Safety Monitors Against Model Compliance"(submitted to JUDGe workshop @ NeurIPS 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑