AI 中文总结
研究人工智能代理与工具交互安全问题,提出ToolGuardian框架,通过预准入审查和任务感知运行时授权保障安全,利用渐进式特征化和基于ASP的声明式策略层,实验表明该框架在审查和授权方面性能良好。
AI 中文摘要
大型语言模型(LLM)代理越来越依赖外部工具,在扩展能力的同时创建了新的安全边界,因为第三方工具在接口层面看似无害,但在实现中可能嵌入不安全行为。现有防御措施存在不足。本文提出ToolGuardian,一个通过预准入审查和任务感知运行时授权来保障代理与工具交互安全的策略驱动框架。它使用渐进式特征化将证据转化为结构化事实。其核心贡献是基于回答集编程(ASP)的声明式策略层,能对能力、效果、任务上下文和组合进行明确推理。通过实验对比,评估了ToolGuardian在多种工具和运行时场景下的表现,结果显示其在审查和运行时授权方面都有较好性能。
英文摘要
LLM agents increasingly rely on external tools, expanding capability while creating a new security boundary: third-party tools may appear benign at the interface level while embedding unsafe behavior in implementation. Existing defenses rely on weak metadata, collapse characterization and policy judgment into a single decision, or use heuristic/LLM enforcement that lacks deterministic, auditable reasoning over task context and multi-tool composition. This paper presents ToolGuardian, a policy-driven framework for securing agent-tool interactions through pre-admission vetting and task-aware runtime authorization. ToolGuardian uses progressive characterization to convert evidence into structured facts: descriptions capture declared intent, system-call traces expose coarse behavior, mock execution reveals observed effects, and source analysis identifies latent behavior. ToolGuardian's core contribution is an Answer Set Programming (ASP)-based declarative policy layer that reasons explicitly over capabilities, effects, task context, and composition. We compare ASP against heuristic and LLM-based policy realizations using identical inputs and output contracts. We evaluate ToolGuardian on 16 MCP-style tools, including 8 malicious variants derived from real open-source tools, and 20 runtime scenarios. For vetting, ASP reaches a deny-class F1 of 0.86 and 88% accuracy using description, syscall, and observed-effect evidence. For runtime authorization, fully specified realizations classify all scenarios correctly, while ablations show that removing compositional and conformance rules substantially degrades performance.