arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21126cs.CR

TraceGrant:一种用于联网LLM智能体任务-效应生命周期的合约治理安全框架

TraceGrant: A Contract-Governed Security Framework for the Task-Effect Lifecycle of Networked LLM Agents

Bohao Liao, Jingchao Wang, Qipeng Song, Jin Cao, Jieling Wang, Boyu Deng

首次发表
浏览论文内容

中文总结 AI 辅助

TraceGrant是通过显式合约治理联网LLM智能体任务-效应生命周期的安全框架,在两类基准攻击案例中未出现攻击成功,且能关联用户意图、执行与任务完成。

中文摘要 AI 辅助

联网大型语言模型(LLM)智能体会从电子邮件、云存储、日历、交易平台及Web服务中检索信息,以完成会产生持久外部效应的多步骤任务。合法执行所需的相同内容中可能也包含间接提示注入,这种注入会重定向工具使用、修改敏感参数或破坏任务完成。现有防御措施主要是约束不受信任的内容或单个工具调用,未能将用户意图、运行时证据、已实现效应与任务完成充分关联。我们提出TraceGrant,这是一种通过显式合约来治理联网LLM智能体任务-效应生命周期的安全框架:执行前,TraceGrant会从可信用户请求中建立任务-效应边界;执行期间,被允许的证据只能实例化合约已确立的权限;执行后,会根据实际工具结果验证任务完成情况。在固定基准设置下,针对949个AgentDojo攻击案例和400个Agent Security Bench攻击案例,TraceGrant在攻击率分别为77.32%和83.00%时未记录到攻击成功,同时保持了实用性。我们还通过白盒防御感知攻击、合约质量分析、阶段消融实验、针对性压力测试及运行时开销测量对TraceGrant进行评估,结果表明TraceGrant提供了一个统一治理层,将可信用户意图、运行时证据、具体工具执行与已验证的任务完成关联起来。

英文摘要

Networked large language model (LLM) agents retrieve information from email, cloud storage, calendars, transaction platforms, and Web services to complete multistep tasks that produce persistent external effects. The same content needed for legitimate execution may also contain indirect prompt injections that redirect tool use, alter sensitive arguments, or disrupt task completion. Existing defenses mainly constrain untrusted content or individual tool calls, leaving user intent, runtime evidence, realized effects, and task completion insufficiently connected. We present TraceGrant, a security framework that governs the task-effect lifecycle of networked LLM agents through an explicit Contract. Before execution, TraceGrant establishes a task-effect boundary from the trusted user request. During execution, admitted evidence can instantiate only authority already established by the Contract. After execution, task completion is verified against actual tool results. Across 949 AgentDojo and 400 Agent Security Bench attack cases under fixed benchmark settings, TraceGrant recorded no attack successes while retaining utility under attack rates of 77.32% and 83.00%, respectively. We further evaluate TraceGrant through white-box defense-aware attacks, Contract quality analysis, stage ablations, targeted stress tests, and runtime overhead measurements. The results show that TraceGrant provides a unified governance layer that connects trusted user intent, runtime evidence, concrete tool execution, and verified task completion.

↑