arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

测量与利用LLM工具调用流水线中的隐式信任

Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines

Murali Ediga, Sudipta Chattopadhyay

arXiv 2609.18217首次发表:更新:

发表机构

University of Missouri-Kansas City(密苏里大学堪萨斯城分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出跨通道碎片化攻击,利用LLM工具调用中多通道共享上下文无特权分离的隐式信任,使完全抵抗单通道注入的模型在双通道下以高达100%比例外泄凭证,且现有安全工具均无法检测。

AI 中文摘要

模型上下文协议(MCP)使大型语言模型(LLM)能够调用外部工具,但每次工具交互都会通过多个输入通道(工具描述、工具结果、采样消息)将模型暴露于攻击者控制的文本中,这些通道共享单一上下文窗口,没有特权分离。在本文中,我们提出了一个框架,通过不同通道发送的各种载荷框架来测量任意LLM的信任概况。在此评估之后,我们设计了跨通道碎片化攻击,将看似良性的载荷分布在两个或三个通道上;没有任何单个通道携带完整的注入,但LLM将碎片编译成凭证外泄。我们在12个前沿模型、三个生产客户端和六个载荷上评估了我们的攻击,总计超过15,000次试验。我们的评估揭示,跨通道攻击是一个未被探索的攻击面:完全抵抗单通道注入(0%合规)的模型在双通道碎片化下以高达100%的比例外泄敏感数据(例如,GPT-4o、Llama 70B、Composer 2、Haiku 4.5)。我们进一步展示了价值对齐的利用,其中工具声明的目的需要攻击者目标的数据,以及通过VS Code的MCP实现注入持久指令的采样系统提示覆盖。最后,我们针对七个第三方MCP安全工具和三种基于提示的防御评估了我们的攻击。所有工具都未能检测到碎片化载荷,而提示防御被证明是模型特定的而非通用的。

英文摘要

The Model Context Protocol (MCP) enables LLMs to invoke external tools, but every tool interaction exposes the model to attacker-controlled text through multiple input channels (tool descriptions, tool results, sampling messages) that share a single context window without privilege separation. In this paper, we present a framework to measure the trust profile of an arbitrary LLM based on a variety of payload framings sent through different channels. Following this assessment, we devise cross-channel fragmentation attacks that distribute seemingly benign payloads across two or three channels; no individual channel carries a complete injection, yet the LLM compiles the fragments into credential exfiltration. We evaluated our attacks across 12 frontier models, three production clients, and six payloads, totalling over 15,000 trials. Our evaluation reveals that cross-channel attacks are an unexplored attack surface: models that fully resist single-channel injection (0% compliance) exfiltrate sensitive data at up to 100% under two-channel fragmentation (e.g., GPT-4o, Llama 70B, Composer 2, Haiku 4.5). We further demonstrate value-aligned exploitation, where a tool's stated purpose requires the data the attacker targets, and a sampling system prompt override that injects persistent instructions via VS Code's MCP implementation. Finally, we evaluated our attacks against seven third-party MCP security tools and three prompt-based defenses. All tools failed to detect fragmented payloads, and prompt defenses proved model-specific rather than universal.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑