用于大语言模型智能体中安全MCP工具使用的混合分析
Hybrid Analysis for Secure MCP Tool Use in LLM Agents
- Zhejiang University(浙江大学)
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究LLM智能体中MCP工具使用的安全问题,提出基于混合分析的MTGuard框架,利用生命周期感知的静态 - 动态协同分析,有效减轻有害工具使用,保持良性任务性能。
AI中文摘要:
大语言模型(LLM)智能体的快速发展使其在各种实际任务中广泛应用。为规范LLM智能体与外部环境的交互,模型上下文协议(MCP)工具成为事实上的标准并被广泛集成。但MCP工具的使用带来新安全风险,此前多数防御方法依赖静态分析,限制了防御效果和鲁棒性。为此提出MTGuard,一个基于混合分析的防御框架,利用生命周期感知的静态 - 动态协同分析保护LLM智能体中MCP工具的使用。大量评估表明MTGuard有效减轻不同LLM智能体中多类有害工具使用,同时保持良性用户任务性能。
英文摘要:
The rapid development of large language model (LLM) agents has enabled their broad adoption across diverse real-world tasks. To standardize interactions between LLM agents and external environments, Model Context Protocol (MCP) tools have emerged as a de facto standard and have been widely integrated into these systems. However, the use of MCP tools also introduces new safety risks, as LLM agents can be induced to perform malicious or unauthorized actions. Although prior work has proposed defenses for securing tool use in LLM agents, most methods rely on static analysis, i.e., inspecting prompts and generated outputs, which limits the defense effectiveness and robustness. To address these limitations, we propose MTGuard, a hybrid analysis-based defense framework designed to safeguard the use of MCP tools in LLM agents by leveraging lifecycle-aware static-dynamic co-analysis. Extensive evaluation demonstrates that MTGuard effectively mitigates multiple categories of harmful tool use across different LLM agents while maintaining performance on benign user tasks.