arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WebMCP-Phalanx:为浏览器集成的大语言模型智能体实施并刻画信任边界

WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents

Lin-Fa Lee, YI-YU Chang, Kuo-Hui Yeh

arXiv 2608.24017首次发表:更新:

发表机构

National Yang Ming Chiao Tung University(国立阳明交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对浏览器集成 LLM 智能体的安全风险,提出 WebMCP-Phalanx 双层架构,验证其可降低攻击成功率、阻断多数提示注入,仅少量工具返回攻击成功,任务效用未受明显影响。

AI 中文摘要

W3C 新兴的 WebMCP 提案使大语言模型(LLM)智能体能够调用网页暴露的工具。然而在多方网页环境中,将智能体执行集成到以同源策略(SOP)为核心的浏览器安全模型中,无法为智能体可访问的工具提供足够的来源和生命周期保障,产生三类风险:主体归属欺骗、不受控的工具生命周期、语义提示注入。我们提出 WebMCP-Phalanx,一种双层智能体运行时架构。其第一层提供浏览器原生信任锚,通过密码学保护的能力凭证将每个工具绑定到其注册主体,并在整个工具生命周期中传播来源标签。第二层将语义检查与特权工具使用分离:无工具调用权限的隔离智能体(Q-LLM)检查工具元数据、输出及网页提供内容的提示注入,验证后的内容再转发给特权智能体(P-LLM)执行,同时 Q-LLM 的内部状态对网页脚本保持隐藏。实证评估显示,浏览器原生所有权机制将撤销和覆盖攻击的成功率从 100% 降至 0%;双层智能体运行时阻断了工具描述中嵌入的全部 80 次提示注入尝试,将工具返回攻击限制在 80 次中成功 2 例;所有实验中,任务效用与无攻击基线无统计差异。不过在白盒自适应攻击者下,基于描述的过滤可通过检查前调用的恶意工具名称绕过,该发现促使需添加调用时序门,将工具调用延迟到所有智能体可见的工具元数据验证完成后。

英文摘要

The emerging W3C WebMCP proposal enables LLM agents to invoke tools exposed by web pages. In multi-party web environments, however, integrating agent execution into a browser security model centered on the Same-Origin Policy (SOP) leaves insufficient provenance and lifecycle guarantees for agent-accessible tools, creating three risks: subject-attribution spoofing, uncontrolled tool lifecycles, and semantic prompt injection. We propose WebMCP-Phalanx, a dual-layer agent runtime architecture. Its first layer provides a browser-native trust anchor that binds each tool to its registering principal through cryptographically protected capability credentials and propagates provenance labels throughout the tool lifecycle. Its second layer separates semantic inspection from privileged tool use. A Quarantine Agent (Q-LLM), without tool invocation authority, inspects tool metadata, outputs, and page-supplied content for prompt injection. Validated content is then forwarded to a Privileged Agent (P-LLM) for execution, while the Q-LLM's internal state remains hidden from page scripts. Empirical evaluation shows that the browser-native ownership mechanism reduces revocation and overwrite attack success from 100\% to 0\%. The dual-agent runtime blocks all 80 prompt-injection attempts embedded in tool descriptions and limits tool-return attacks to 2 successful cases out of 80. Across experiments, task utility remains statistically indistinguishable from the no-attack baseline. Under a white-box adaptive attacker, however, description-based filtering can be bypassed through malicious tool names invoked before inspection. This finding motivates a call-timing gate that delays tool invocation until all agent-visible tool metadata has been validated.

Comments8 pages, 1 figure, AAAI2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑