arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03470cs.CR

CorrectGuard:黑盒安全护栏的免观察正确性估计

CorrectGuard: Eyes-Off Correctness Estimation for Black-Box Security Guardrails

Adam Faulkner, Nil-Jana Akpinar, Matthew Dressman

首次发表
浏览论文内容

中文总结 AI 辅助

CorrectGuard提出免观察正确性估计框架,用独立模型仅凭可见数据预测黑盒护栏决策正确性,在13个数据集上显著提升错误识别,并支持决策排序与弃权(不执行)。

中文摘要 AI 辅助

AI服务越来越依赖黑盒安全护栏,然而保护隐私的模型审计机制往往无法衡量这些系统在两种设置下的表现:一种是人类免观察的生产环境,禁止人工检查用户输入;另一种是机器免观察的环境,禁止模型检查此类输入。我们提出了CorrectGuard,一种适用于这两种设置的免观察正确性估计框架,它涉及一个独立的基于模型的评估器,仅使用带标签的可见数据,在不访问护栏内部机制的情况下,预测护栏对人和机器不可访问输入的决策是否正确。我们在13个安全与安全数据集上(涵盖有害内容、越狱、提示注入和信息提取)以及统一视为黑盒的开放权重护栏上,采用留一数据集评估方法,评估了基于上下文学习、嵌入和微调的正确性模型。在人类和机器免观察设置(后者使用保护隐私的输入指纹识别实现)中,基于上下文学习的正确性分类器显著改善了跨护栏的错误识别,宏准确率最高提升25个百分点;基于微调的方法也提供了近15个百分点的提升,尽管性能在不同护栏和留出数据集上差异显著。正确性分数还支持护栏决策排序和弃权(不执行):在3个护栏上,最佳正确性排序将AURC从无排序基线的0.33-0.44降至0.17-0.22,而最佳操作点在15%的观测风险下保留了37.5-52.0%的护栏决策。这些结果表明,外部正确性模型无需特权访问护栏即可暴露系统性失败并支持护栏决策弃权(不执行)。

英文摘要

AI services increasingly rely on black-box security guardrails, yet privacy-preserving model auditing regimes often cannot measure how well these systems perform in both a human eyes-off production setting, which disallows human inspection of user input, and a machine eyes-off setting, which disallows model inspection of such input. We introduce CorrectGuard, an eyes-off correctness estimation framework for both settings, which involves an independent model-based evaluator predicting whether guardrail decisions on human- and machine-inaccessible inputs are correct using only labeled eyes-on data and without access to the guardrail's internals. We evaluate in-context learning, embedding, and finetuning-based correctness models under leave-one-dataset-out evaluation across 13 safety and security datasets spanning harmful content, jailbreaks, prompt injection, and extraction, and across open-weight guardrails treated uniformly as black boxes. Across both human and machine eyes-off settings (the latter implemented using privacy-preserving fingerprinting of inputs), in-context-learning-based correctness classifiers substantially improve error identification across guardrails, achieving up to a 25 percentage-point increase in macro accuracy, as do finetuning-based approaches which provide a nearly 15-point boost, although performance varies sharply across guardrails and held-out datasets. Correctness scores also support guardrail decision ranking and abstention: across 3 guardrails, the best correctness rankings reduce AURC from unranked baselines of 0.33-0.44 to 0.17-0.22, while the best operating points retain 37.5-52.0% of guardrail decisions at 15% observed risk. These results show that external correctness models can expose systematic failures and support guardrail decision abstention without privileged access to the guardrail.

发表机构

  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

↑