AgentBoundary:工具使用型LLM智能体的安全性反事实评估
AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents
浏览论文内容
中文总结 AI 辅助
针对工具使用型LLM智能体安全评估中过度拒绝与任务失败混淆的问题,提出四路反事实框架AgentBound,通过区分表面风险与行动许可性,诊断过度拒绝和不安全合规,并训练校准模块提升授权任务完成率。
中文摘要 AI 辅助
大型语言模型(LLM)在对话场景中的安全对齐主要围绕是否回答或拒绝请求来构建。然而,在智能体场景中,同样的模型必须在执行过程中出现权限关键证据时决定是否采取行动。这带来了一个独特的挑战:表面风险、行动许可性和任务能力很容易混淆,使得智能体的过度拒绝难以与普通任务失败区分开来。为解决这一问题,我们引入了AgentBound,这是首个用于工具使用型智能体安全性的四路反事实生成与评估框架。AgentBound通过独立变化表面风险和行动许可性来转换相同的可执行工作流,从而能够对看似有风险但已授权的任务和看似常规但未授权的任务进行受控比较。这些比较在控制任务能力的同时,共同诊断过度拒绝和不安全合规问题。我们将AgentBound实例化为一个经过人工验证的4000任务评估套件,包含基于轨迹和基于事后状态的判断。在17种模型和框架配置中,高安全性常常与较差的授权任务完成率并存:GPT-5.5阻止了99.5%的看似常规的未授权操作,却仅完成了28.7%的看似有风险的授权任务。我们进一步训练了一个轻量级运行时校准模块,在10种评估配置中平均将授权任务完成率提高了18.2%,同时平均将不安全操作阻止率提高了5.4%。这些结果表明,有效的智能体对齐要求行动决策跟踪与权限相关的执行证据,而不仅仅是依赖拒绝强度。
英文摘要
Safety alignment for large language models (LLMs) in conversational settings is largely framed around whether to answer or refuse a request. In agentic settings, however, the same models must decide whether to act as permission-critical evidence emerges during execution. This creates a distinct challenge: apparent risk, action permissibility, and task competence are easily confounded, making agentic over-refusal difficult to distinguish from ordinary task failure. To address this, we introduce AgentBound, the first four-way counterfactual generation-and-evaluation framework for tool-using agent safety. AgentBound transforms the same executable workflow by independently varying apparent risk and action permissibility, enabling controlled comparisons of risky-looking but authorized tasks and routine-looking but unauthorized tasks. These comparisons jointly diagnose over-refusal and unsafe compliance while controlling for task competence. We instantiate AgentBound as a human-validated 4,000-task evaluation suite with trajectory-based and post-state-based judgments. Across 17 model and harness configurations, high safety frequently coexists with poor authorized-task completion: GPT-5.5 blocks 99.5\% of routine-looking unauthorized actions yet completes only 28.7\% of risky-looking authorized tasks. We further train a lightweight runtime calibration module that improves authorized-task completion by 18.2\% on average across 10 evaluated configurations, while improving unsafe-action blocking by 5.4\% on average. These show that effective agentic alignment requires action decisions to track permission-relevant execution evidence, rather than refusal strength alone.
发表机构
- Peking University(北京大学)
- Beijing Academy of Artificial Intelligence(北京人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。