arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23763cs.CRcs.AI

TrustShiftProbe:对MCP服务器上的阶段性信任攻击进行表征、基准测试与防御

TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers

Mehrdad Rostamzadeh, Sidhant Narula, Mohammad Ghasemigol, Daniel Takabi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对MCP服务器的阶段性信任攻击提出TrustShiftProbe框架,建立威胁模型、攻击引擎与防御机制SHIELD,使攻击成功率从69.5%降至42.7%

中文摘要 AI 辅助

模型上下文协议(Model Context Protocol, MCP)已成为连接大语言模型智能体与外部工具后端的标准层。这种开放性引入了一种严重的服务器端威胁,我们称之为信任转移(TrustShift):被入侵的MCP服务器在初始条件阶段表现正常,建立操作依赖并抑制智能体的怀疑态度,一旦达到交互阈值就切换到对抗性有效载荷。这种规避是时间性的,而非句法性的:在部署时表现正常,服务器的叛变对部署前静态分析不可见,静态分析仅能看到诚实阶段。切换后的有效载荷范围从明显的结构违规到符合模式的操纵,后者保留了外部协议合规性以规避运行时中间件过滤器。关键在于,信任转移源于服务器控制的工具通道,而非用户提示(与间接提示注入不同)或传输层(与中间人攻击不同):攻击者本身是受信任的服务器端点。我们推出TrustShiftProbe,这是一个评估与防御框架,包含四项贡献:(1)将智能体-服务器生命周期建模为有状态的时间威胁模型,即良性条件阶段后在信任 horizon 处发生对抗性叛变;(2)一种与语言无关的攻击引擎,将每种变体实例化为四个生产领域中被入侵的MCP服务器;(3)SHIELD,一种在MCP传输边界处的多层、零预言机运行时防御,针对在干净信任窗口期间学习到的行为基线审计服务器有效载荷;(4)涵盖三种执行机制(结构违规、语义破坏、范围扩展)和三种对抗目标(破坏、数据 exfiltration 及二者结合)的九种信任转移变体分类。在前沿专有与开源模型上,信任转移攻击达到69.5%的平均攻击成功率,SHIELD将其缓解至42.7%。

英文摘要

The Model Context Protocol (MCP) has emerged as the standard layer connecting Large Language Model agents to external tool backends. This openness introduces a severe server-side threat we term TrustShift: a compromised MCP server behaves benignly during an initial conditioning phase, building operational reliance and suppressing agent skepticism, before switching to an adversarial payload once an interaction threshold is reached. The evasion is temporal, not syntactic: benign at deploy time, the server's defection is invisible to predeployment static analysis, which sees only the honest phase. Switched payloads range from overt structural violations to schema-valid manipulations, the latter preserving outer protocol compliance to evade runtime middleware filters. Crucially, TrustShift originates in the server-controlled tool channel, not user prompts (unlike indirect prompt injection) or the transport (unlike man-in-the-middle): the adversary is the trusted server endpoint itself. We introduce TrustShiftProbe, an evaluation and defense framework with four contributions: (1) a stateful temporal threat model of the agent-server lifecycle as a benign conditioning phase followed by an adversarial defection at a trust horizon; (2) a language-agnostic attack engine that instantiates each variant as a compromised MCP server across four production domains; (3) SHIELD, a multi-tier, zero-oracle runtime defense at the MCP transport boundary that audits server payloads against behavioral baselines learned during clean trust windows; and (4) a taxonomy of nine TrustShift variants spanning three execution mechanisms (structural violation, semantic corruption, scope expansion) and three adversarial objectives (disruption, exfiltration, and their combination). Across frontier proprietary and open-weight models, TrustShift attacks achieve a 69.5% mean attack success rate that SHIELD mitigates to 42.7%.

发表机构

  • Old Dominion University(奥多明尼昂大学)

机构由 AI 辅助整理,请以论文原文为准。

↑