arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

持久可计费状态:工具调用LLM智能体中的拒绝钱包攻击与防御

Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents

Jinqian Zhang, Haojun Xia, Shujiang Wu, Jingkun Yue, Xia Zhang, Zhangpei Cheng, Bibo Tu

arXiv 2609.28585首次发表:更新:

发表机构

Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences; Beihang University; Beijing University of Posts and Telecommunications(中国科学院信息工程研究所; 中国科学院大学网络空间安全学院; 北京航空航天大学; 北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次系统分析工具调用LLM智能体的持久可计费状态,揭示拒绝钱包攻击,提出宿主侧不变量与历史转换防御,并验证其有效性。

AI 中文摘要

多步骤工具调用LLM智能体依赖宿主运行时在轮次之间保持状态。当运行时将外部工具返回值携带到后续模型输入中时,提供商会再次计量。一个被承认的恶意或受损工具因此可以将不可信数据转化为重复的受害者计费处理,而无需受害者凭据或本地运行时权限。我们将保留的内容称为持久可计费状态,并将宿主关于其是否以及如何进入后续可计费上下文的决策形式化为持久可计费状态边界。我们首次对这一准入后生命周期进行了系统性安全研究。我们推导出六种拒绝钱包攻击向量,并构建了DOW-BENCH,一个在六个模型家族上评估的端到端测试平台。在243次执行中,使用遥测显示,每次会话的最大累计输入达到会话首次调用输入的14,293倍。受控的历史策略重跑隔离了原始保留的贡献:保留原始历史使平均有效会话成本增加21.2%至35.9%。压缩在10/12和11/12的历史依赖任务上成功,而删除在每个提供商的2/12上成功。为了管理这一边界,我们将确定性历史转换与四个宿主侧不变量相结合,这些不变量在重新摄入前限制提示质量、上下文增长、递归机会和累计支出。内核包含123次评估重放语料库中的每一次重复攻击。在24个Mistral Small 4工作流中,进度授权策略实现了22/24的预言机验证任务成功,且无完成前中断,而固定上限下为13/24。在3,830个扫描的MCP服务器和传输仓库中,只有71个暴露了任何代码可见的安全防护代理,且没有一个覆盖全部四个防护家族。这些结果确立了持久可计费状态作为一级安全对象,以及重新摄入前作为其宿主拥有的控制点。

英文摘要

Multi-step tool-calling LLM agents rely on host runtimes to preserve state across turns. When a runtime carries an external tool return into later model inputs, providers meter it again. An admitted malicious or compromised tool can thereby convert untrusted data into recurring victim-billed processing without victim credentials or local runtime privilege. We call retained content persistent billable state and formalize the host's decision over whether and how it enters later billable context as the persistent billable-state boundary. We present the first systematic security study of this post-admission lifecycle. We derive six denial-of-wallet attack vectors and build DOW-BENCH, an end-to-end harness evaluated across six model families. Across 243 executions, usage telemetry shows that the maximum per-session cumulative input reaches 14,293x the session's first-call input. Controlled history-policy reruns isolate raw retention's contribution: retaining raw history increases mean effective session cost by 21.2-35.9%. Compression succeeds on 10/12 and 11/12 history-dependent tasks, versus 2/12 under deletion for each provider. To govern this boundary, we combine deterministic history transformation with four host-side invariants that bound prompt mass, context growth, recursive opportunity, and cumulative spend before reingestion. The kernel contains every recurring attack in the 123-evaluation replay corpus. Across 24 Mistral Small 4 workflows, a progress-authorized policy achieves 22/24 oracle-verified task successes with no pre-completion interruptions, versus 13/24 under a fixed cap. Only 71 of 3,830 scanned MCP server and transport repositories expose any code-visible safeguard proxy, and none cover all four safeguard families. These results establish persistent billable state as a first-class security object and pre-reingestion as its host-owned control point.

Comments22 pages, 14 figures, 13 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑