arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39450cs.CRcs.AI

ActionGuard:受污染技能下的工具调用授权

ActionGuard: Tool Call Authorization under Poisoned Skills

  • Korea University(高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

Jihun Han, Yejin Jang, Byung Il Kwak, Mee Lan Han

AI总结:

针对技能注入导致工具调用被劫持的问题,提出ActionGuard,通过分离生成与授权上下文、基于用户意图和运行时证据进行执行前审查,显著降低攻击成功率并保持任务完成率。

AI中文摘要:

基于LLM的智能体通过第三方技能扩展其能力,这些技能提供特定任务的指令、脚本和工具使用流程。然而,恶意指令被插入到原本良性的技能中,可能导致良性用户请求触发危险的工具调用,包括数据外泄、文件删除或未经授权的代码执行。本文提出了ActionGuard,它在工具调用即将执行之前对其进行检查。ActionGuard将目标智能体的动作生成上下文与安全防护的授权上下文分离。目标智能体可以使用原始技能进行规划,但审查者不会接收到可能受污染的原始技能文本。相反,它通过平衡的技能概况、当前和最近的工具调用以及本地脚本内容,判断每个动作是否由可信的用户请求证明合理。ActionGuard在OpenClaw的“工具调用前”阶段拦截每个工具调用,并在故障关闭策略下强制执行审查者的允许或拒绝决定。我们在SKILL-INJECT设置下,针对Dynamic Guardian和SkillGuard,使用三个开源和两个商业审查者模型,对139个上下文注入和180个明显注入进行了评估。每个条件重复三次,并使用攻击成功率(ASR)和任务成功率(TSR)进行评估。总体而言,与现有防护相比,ActionGuard将ASR相对降低了35.54%至46.11%,与无防护相比降低了70.44%,同时保持了较高的良性任务完成率。这些结果表明,基于可信用户意图和运行时证据的执行边界授权可以限制由技能注入引起的未经授权的工具调用。

英文摘要:

LLM-based agents extend their capabilities through third-party skills that provide task-specific instructions, scripts, and tool-use procedures. However, malicious instructions inserted into an otherwise benign skill can cause a benign user request to trigger dangerous Tool Calls, including data exfiltration, file deletion, or unauthorized code execution. This paper presents ActionGuard, which inspects skill-influenced Tool Calls immediately before execution. ActionGuard separates the target agent's action-generation context from the safeguard's authorization context. The target agent may use the original skill for planning, but the Reviewer does not receive the potentially poisoned raw skill text. Instead, it determines whether each action is justified by the trusted user request using a balanced skill profile, current and recent Tool Calls, and local script contents. ActionGuard intercepts each Tool Call at OpenClaw's before-tool-call stage and enforces the Reviewer's ALLOW or DENY decision under a fail-closed policy. We evaluate ActionGuard on 139 contextual and 180 obvious injections in a SKILL-INJECT-based setting against Dynamic Guardian and SkillGuard, using three open-source and two commercial Reviewer models. Each condition is repeated three times and evaluated using Attack Success Rate (ASR) and Task Success Rate (TSR). Overall, ActionGuard reduced ASR by 35.54 to 46.11 percent relative to existing safeguards and by 70.44 percent relative to No Safeguard, while maintaining high benign-task completion. These results show that execution-boundary authorization grounded in trusted user intent and runtime evidence can restrict unauthorized Tool Calls induced by skill injection.

↑