arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BLINDSPOT:长时程工具使用智能体的安全性与拒绝校准基准

BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

Sadia Asif, Mohammad Mohammadi Amiri, Momin Abbas, Tejaswini Pedapati, Prasanna Sattigeri

arXiv 2609.16305首次发表:更新:

发表机构

Rensselaer Polytechnic Institute; IBM Research(伦斯勒理工学院; IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长时程工具使用智能体,提出轨迹级安全校准基准Blindspot,通过自适应对抗交互评估多轮安全,揭示模型间显著差异及延迟故障现象。

AI 中文摘要

大型语言模型(LLM)智能体越来越多地参与涉及工具使用、持久状态、不断演化的授权以及外部环境反馈的长时程交互。在此类场景中,安全故障可能仅在多轮交互后才显现,然而现有评估往往将智能体行为简化为任务或攻击成功率,从而掩盖了智能体在交互演化过程中是执行、拒绝还是保持适当校准。我们引入了Blindspot,一个用于长时程工具使用智能体轨迹级安全校准的基准。Blindspot通过自适应对抗交互、有状态工具执行和基于执行的裁决来评估完整的用户-智能体-环境轨迹。其当前版本包含7个领域的22个攻击家族和35个场景,生成了超过2,500条长时程轨迹,平均交互长度为14.7轮。每条轨迹被赋予五种结果之一:安全完成、正确拒绝、不安全完成、过度拒绝或不确定。与固定攻击数据集不同,Blindspot是一个可扩展的实时模拟框架,其中攻击、场景、工具、策略、领域和智能体配置均可添加,而无需重新设计评估流程。我们使用8个指标评估了13个专有和开放权重LLM,涵盖不安全完成、适当拒绝、良性效用、过度拒绝、重复运行鲁棒性和拒绝后失败。初步结果揭示了各模型在安全-效用校准上的显著差异,并表明故障可能仅在若干初始安全交互步骤之后才出现。这些发现促使我们将智能体安全视为轨迹级属性,而非单轮或二元成功标准。

英文摘要

Large language model (LLM) agents increasingly operate over long-horizon interactions involving tool use, persistent state, evolving authorization, and external environment feedback. In such settings, safety failures may emerge only after multiple turns, yet existing evaluations often reduce agent behavior to task or attack success, obscuring whether an agent acts, refuses, or remains appropriately calibrated as the interaction evolves. We introduce Blindspot, a benchmark for trajectory-level safety calibration of long-horizon tool-using agents. Blindspot evaluates complete user-agent-environment trajectories through adaptive adversarial interaction, stateful tool execution, and execution-grounded adjudication. Its current instantiation contains 22 attack families and 35 scenarios across seven domains, yielding more than 2,500 long-horizon trajectories with an average interaction length of 14.7 turns. Each trajectory is assigned one of five outcomes: Safe Completion, Correct Refusal, Unsafe Completion, Over-Refusal, or Indeterminate. Unlike fixed attack datasets, Blindspot is an extensible live-simulation framework in which attacks, scenarios, tools, policies, domains, and agent configurations can be added without redesigning the evaluation pipeline. We evaluate 13 proprietary and open-weight LLMs using eight metrics covering unsafe completion, appropriate refusal, benign utility, over-refusal, repeated-run robustness, and post-refusal failure. Preliminary results reveal substantial differences in safety-utility calibration across models and show that failures can emerge only after several initially safe interaction steps. These findings motivate treating agent safety as a trajectory-level property rather than a single-turn or binary success criterion.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑