arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越直接访问:大语言模型智能体中的资源劫持

Beyond Direct Access: Resource Hijacking in LLM Agents

Puyu Zeng, Mingang Chen, Zheli Liu, Qibing Ren

arXiv 2608.15108首次发表:更新:

发表机构

College of Cryptology and Cyber Science, Nankai University; Shanghai Jiao Tong University(南开大学密码与网络科学学院; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究首次识别智能体资源劫持这一安全盲区,推出ResourceHijackBench及自动化生成流程,实验显示无防御时OpenClaw平均攻击成功率达84.06%,现有防御仍不足。

AI 中文摘要

大语言模型智能体正日益连接到计算基础设施、凭证、使用预算、身份、私有知识、通信渠道以及组织工作流等高价值资源。现有智能体安全研究主要针对指令、数据和工具行为的攻击展开,而智能体可访问的高价值资源作为直接攻击目标受到的关注则少得多。我们首次识别并系统研究了智能体资源劫持这一安全盲区:攻击者诱导智能体为自身目标调用、消耗、转移或控制高价值资源,且无需直接获取这些资源或其凭证。为研究该威胁,我们推出了ResourceHijackBench及用于生成资源劫持案例的自动化流程。我们将智能体的高价值资源分为六大类,构建了包含900个攻击提示的300个攻击场景。每个案例在隔离的本地环境中运行,记录实际资源使用情况,从而可仅从智能体行为而非文本响应来评估攻击。在无额外防御措施的情况下,OpenClaw的平均攻击成功率达84.06%;该攻击在不同模型后端均有效,平均成功率介于69.98%至89.58%之间。现有防御措施可降低部分风险,但评估中最强的防御仍使平均攻击成功率达55.11%。这些结果表明,智能体可访问的高价值资源构成了此前被忽视的重要攻击面,且当前智能体防御措施不足以保护其免受资源劫持。

英文摘要

Large language model agents are increasingly connected to high-value resources, including external APIs, GPUs and servers, and workflows such as deployment and approval. Existing agent security research mainly focuses on attacks against information and agent behavior, while high-value resources have received less attention as attack targets themselves. To our knowledge, we are the first to identify and systematically study agent resource hijacking, in which attackers induce agents to use high-value resources for their own goals without directly obtaining those resources or their credentials. We introduce ResourceHijackBench, an executable benchmark and automated case-generation pipeline that covers six categories of high-value resources. It contains 300 attack scenarios and 900 attack prompts, and each case runs in an isolated local environment that records actual resource use. Resource hijacking remains effective across four model backends, with average attack success rates of 70.0% to 89.6%, and it also appears across two agent harnesses, with average ASRs of 84.1% on OpenClaw and 72.3% on Codex. In paired comparisons, resource hijacking achieves ASRs 62.4 to 84.0 percentage points higher than direct resource acquisition across the four model backends. Real-world experiments also demonstrate the practical effectiveness of resource hijacking. Among the three existing defenses we evaluate, the lowest average ASR is still 55.1%. We further propose ResGate, a pre-execution resource authorization defense that combines model-based resource-use extraction with deterministic policy enforcement based on trusted requester identity metadata, reducing the average ASR on OpenClaw to 23.6%. These results show that preventing direct access alone is not enough to protect resources that agents can still use on an attacker's behalf, and that explicit resource authorization can help reduce this risk.

Comments21 pages, 5 figures, 20 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑