发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于溯源分析的ProvenanceGuard框架,通过多阶段流水线检测工具调用前的三种失调类型,在Agent-SafetyBench和WorkBench上显著降低失调错误率并减少不必要的干预。
AI 中文摘要
随着LLM智能体获得越来越强大的工具访问权限,确保其行为与用户意图一致变得至关重要。当智能体提出的工具调用偏离用户意图——这种现象称为失调——可能导致难以撤销的有害后果。现有的运行时防护依赖于LLM作为评判者的范式,缺乏系统性的推理框架来评估对齐,常常产生不一致或难以审计的判断。受溯源分析启发,我们提出了一个基于溯源的概念框架,将失调检测形式化为判断提议的工具调用是否得到智能体上下文中可追溯证据的支持。基于此框架,我们提出了ProvenanceGuard,一个多阶段流水线,在所选工具执行之前分析智能体行为的三种失调类型,并且仅当该行为被认为与用户输入查询对齐时才允许执行。我们在两个不同的基准测试Agent-SafetyBench和WorkBench上,使用10个骨干LLM评估了我们提出的方法。与LLM作为评判者的基线相比,ProvenanceGuard在Agent-SafetyBench上将失调轨迹的错误率从42.9%降低到1.8%,在WorkBench上从32.1%降低到17.3%,同时将任务成功轨迹上的干预负担从30.5%降低到12.8%,并且在对齐轨迹上未引入统计上显著的不必要干预增加。这些结果表明,结构化的、基于溯源的推理为保护LLM智能体免受失调影响提供了有效且实用的基础。
英文摘要
As LLM agents gain increasing access to powerful tools, ensuring that their actions align with the user's intent becomes critical. When an agent's proposed action deviates from that intent---a phenomenon called misalignment---it may cause harm that is difficult to undo. Existing runtime guardrails rely on an LLM-as-a-judge paradigm that lacks a systematic framework for reasoning about alignment, often producing inconsistent or difficult-to-audit judgments. Motivated by provenance analysis, we propose a conceptual framework that formalizes misalignment detection as determining whether a proposed tool call is supported by traceable evidence in the agent's context. Based on this framework, we build ProvenanceGuard, a multi-stage pipeline that analyzes the agent's action for three types of misalignment before its execution and only allows aligned actions. We evaluated ProvenanceGuard on AgentSafetyBench and WorkBench, across 11 backbone LLMs. Compared to the LLM-as-a-judge baseline, ProvenanceGuard reduces error rate on misaligned traces from 44.3% to 2.1% on Agent-SafetyBench and from 32.4% to 18.7% on WorkBench, while reducing interventions on task-successful traces from 31.2% to 13.0% and introducing no statistically significant increase in unnecessary interventions on aligned traces. These results demonstrate that structured, provenance-based reasoning provides an effective and practical foundation for safeguarding LLM agents from misalignment.