发表机构
Tsinghua University; College of AI, Tsinghua University(清华大学; 清华大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究指出仅靠防护模型识别智能体的被禁止行动无法保障安全,提出需识别未履行的义务,构建首个义务识别基准ObligationBench及模型ObligationGuard,显著提升了义务识别性能。
AI 中文摘要
防护模型正越来越多地用于保护基于大语言模型(LLM)的智能体,主要通过识别智能体被禁止执行的行动来实现。然而,仅识别被禁止的行动不足以确保智能体的安全性。在本文中,我们认为智能体的安全性还取决于识别所需但未执行的安全关键行动,我们将这些行动称为义务。我们在一个流行的安全评估基准上进行的初步研究显示,GLM-5.3的轨迹中有56.92%包含未履行的义务,而包含被禁止行动的比例仅为30.00%。这一发现表明,未履行的义务是一个此前被忽视的主要安全风险来源。然而,据我们所知,目前还没有现有的基准来评估防护模型是否能够识别这些义务。为了填补这一空白,我们推出了ObligationBench,这是首个用于评估义务识别能力的基准,包含240个经专家验证的轨迹,涵盖问题解决、功能开发和终端操作场景。我们对14个代表性模型的评估显示出存在显著的局限性:最高的召回率和精确匹配率分别仅为48.97%和10.00%。为解决这些局限性,我们使用40000个合成训练示例开发了ObligationGuard。ObligationGuard的召回率达到57.52%,精确匹配率为21.67%,在这两个指标上均超过了所有被评估的模型。我们呼吁社区将义务识别纳入未来防护模型的设计和评估中,以提升智能体的安全性。
英文摘要
Guard models are increasingly used to safeguard LLM-based agents, primarily by identifying actions that agents are forbidden to perform. However, identifying forbidden actions alone is insufficient to ensure agent safety. In this paper, we argue that agent safety also depends on identifying required yet unperformed safety-critical actions, which we call obligations. Our preliminary study on a popular benchmark for evaluating safety shows that 56.92% of GLM-5.3 trajectories contain unfulfilled obligations, compared with only 30.00% containing forbidden actions. This finding reveals unfulfilled obligations as a major and previously overlooked source of safety risk. However, to our knowledge, no existing benchmark evaluates whether guard models can identify these obligations. To close this gap, we introduce ObligationBench, the first benchmark for evaluating the capability of obligation identification, comprising 240 expert-validated trajectories covering issue resolution, feature development, and terminal operations. Our evaluation of 14 representative models reveals substantial limitations: the highest recall and exact-match rate are only 48.97% and 10.00%, respectively. To address these limitations, we develop ObligationGuard using 40,000 synthetic training examples. ObligationGuard achieves 57.52% recall and an exact-match rate of 21.67%, surpassing all evaluated models on both metrics. We call on the community to incorporate obligation identification into the design and evaluation of future guard models to improve agent safety.
Comments21 pages. Code, data, and supplementary materials: https://github.com/THU-Agent/ObligationGuard