AI 中文总结
本研究构建了首个针对智能体防护栏过度安全问题的基准Cautious Bench,发现防护栏存在名称迷信效应,会因对象名称而非授权上下文过度拒绝合法操作。
AI 中文摘要
智能体防护栏是在大型语言模型(LLM)执行每个操作前对其进行批准或拒绝的检查机制,有时会拒绝真正安全的请求,这种过度安全会在防护栏拒绝授权任务时阻碍部署。评估过度安全的难度在于,授权操作与未授权操作在边界处相似,安全与不安全的标签是授权策略的选择,而非操作本身固定的属性,因此需要一个目前不存在的、能映射理想防护栏决策边界的基准。从真实数据中收集此类基准不切实际,因为边界案例难以收集且其标签难以验证。为此,我们构建了Cautious Bench,这是首个针对智能体防护栏、以过度安全为构建目标的基准,它结合每个样本及其对应的授权政策进行设计;构建时的闸门会重新推导每个示例以进行验证,因此每个标签都是政策的机械结果,而非注释者的逐样本判断,可作为研究人员衡量真实防护栏的参考。该基准生成了756个可判定的良性/孪生对,每个对属于三种对象名称类型(共2268个测量对),另有40个不可判定对单独报告。对来自五种设计的六个防护栏进行测量后,我们发现了名称迷信效应:在对比实验中,仅对象名称不同的情况下,每个防护栏在恐怖外观的对象名称下,比在良性名称下更频繁地过度拒绝授权操作,这种偏差源于防护栏读取表面标签而非授权上下文。
英文摘要
Agent guardrails are checks that approve or refuse each action before an LLM executes it. Sometimes they refuse requests that are genuinely safe. This over-safety blocks deployment when a guardrail refuses an authorized task. Evaluating over-safety is hard: at the boundary an authorized action resembles an unauthorized one, and the safe-versus-unsafe label is a choice of authorization policy, not fixed by the action alone. We argue it therefore requires a benchmark that does not yet exist, one that maps the decision boundary of an ideal guardrail. Harvesting such a benchmark from real data is impractical: boundary cases are hard to collect, their labels hard to verify. The gap is real, so we construct Cautious Bench, the first benchmark to make over-safety the construct for agent guardrails; it codesigns each sample and its label with a stated authorization policy. A build-time gate re-derives every example to certify it, so each label is a mechanical consequence of the policy rather than an annotator's per-sample verdict, a reference against which researchers can measure real guardrails. The benchmark renders 756 Decidable benign/twin pairs, each under three object-name types (2,268 measured pairs), and 40 Undecidable pairs reported separately. Measuring six guardrails from five designs, we find a name-superstition effect: each over-refuses an authorized action more often under a scary-looking object name than a benign one. Since only the object name varies in the aforementioned contrast experiments, the deviation is the name's doing: the guardrails read the surface label, not the authorization context.