评估系统一模型在智能体安全决策中的表现:可靠性、校准与选择性自动化
Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation
- TraceStone
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究评估四种系统一模型在智能体安全决策中的可靠性,发现其平均校准良好但特定攻击上存在系统性失败,且严格漏报限制下自动化程度低,需综合评估准确性、校准与决策策略。
AI中文摘要:
基于模型的评判器通过检测提示注入、评估交互风险以及筛选有害请求来支持智能体安全。系统一模型提供带有概率的类型化决策,软件可利用这些概率来允许、阻止或升级输入,但这些概率是否支持可靠的自动化安全决策仍不明确。我们评估了Jev、Laya、Decider和Bespoke Nimble,并与专门的分类器和语言模型评判器进行比较,考察了决策准确性、概率校准和选择性自动化。我们发现了三个主要结果。(1)强大的整体性能和良好的平均校准可能掩盖特定攻击组上的系统性失败,包括模型自信地分类为安全的攻击。(2)在对漏报攻击的严格限制下,评估的策略很少自动允许输入,而选择单独的允许和阻止阈值主要通过阻止更多输入来增加自动化。即使在验证期间满足错误限制的策略,在未见输入上也可能超出这些限制。(3)第二个评判器可以检测到一些漏报攻击,但它也可能拒绝更多良性输入,并重复第一个模型的高置信度错误。这些发现表明,模型准确性、概率校准以及由此产生的决策策略的行为必须一起评估。
英文摘要:
Agentic software connects language models to tools that modify files and interact with external services. Developers use model-based security checks to screen external content and user requests before agents act. System One models answer developer-defined questions with probabilities over predefined answers such as safe or unsafe, but classification accuracy alone does not establish whether these probabilities support reliable automation. We evaluate Jev, Laya, Decider, and Bespoke Nimble against specialized classifiers and language-model judges across prompt-injection detection, interaction-risk judgment, and harmful-request screening, examining decision accuracy, calibration, and selective automation. (1) High overall accuracy and low average calibration error can conceal attacks classified as safe with high confidence within particular groups. Developers should test candidate models on the intended security task and examine errors within relevant attack groups. (2) Policies selected under strict miss limits allow few test inputs automatically, and choosing allow and block thresholds separately increases automation mainly through additional blocks. Developers should verify limits on missed unsafe inputs and blocked benign inputs on independent data, and report the allowed and blocked fractions separately. (3) Comparing predictions on the same inputs shows that a language-model judge can detect unsafe inputs missed by a System One model, but can also repeat the System One model's confident mistakes and falsely flag benign inputs. Developers should test which missed unsafe inputs their review rule forwards, then measure the review model's misses and benign false alarms on those inputs.