发表机构
School of Cyber Science and Technology, Shandong University; State Key Laboratory of Cryptography and Digital Economy Security, Shandong University; Shandong Key Laboratory of Artificial Intelligence Security, Shandong University(山东大学网络空间安全学院; 山东大学密码与数字经济安全国家重点实验室; 山东大学人工智能安全山东省重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对LLM智能体中工具介导的知识提取风险,提出仅基于查询的提取攻击ToolSiphon,通过工具对比分析与证据链式反馈两个信号,在多工具、多数据集上实现了较高的源记录恢复率,且对防御措施和现实平台有效。
AI 中文摘要
大语言模型(LLM)智能体通常会使用基于知识的工具,并通过工具调用访问其底层文件、数据库和搜索索引。这种集成提升了智能体提供特定领域服务的能力,但也引入了工具介导的知识提取风险:为合法响应而暴露给智能体的源内容可能会从其输出中逐步恢复,从而能够重建目标工具背后的知识源。本文系统研究了这一风险,并确定了工具调用带来的两个挑战:工具选择不确定性,即智能体可能会调用竞争工具而非目标工具;工具参数压缩,即智能体生成工具参数时可能会丢失细粒度查询信息。为应对这些挑战,我们提出了ToolSiphon,一种仅基于查询的提取攻击,它引入了两个互补信号:通过工具对比分析实现的目标区分信号,以引导查询指向目标工具;以及通过证据链式反馈实现的基于响应的事实信号,以缓解参数压缩并逐步扩大提取覆盖范围。在三类基于知识的工具和六个特定领域数据集上,当存在非目标工具的粗粒度信息时,ToolSiphon平均可恢复74.3%的源记录,其中文本恢复率为83.2%,语义相似度为90.2%;即使没有此类信息,它也能恢复66.3%的源记录。ToolSiphon对代表性防御措施以及三个现实世界智能体平台仍保持有效。
英文摘要
LLM agents commonly use knowledge-based tools and access their underlying files, databases, and search indexes through tool invocation. This integration improves agents' ability to provide domain-specific services but also introduces the risk of tool-mediated knowledge extraction: source content exposed to an agent for legitimate responses may be progressively recovered from its outputs, enabling reconstruction of the knowledge source behind a target tool. This paper systematically investigates this risk and identifies two challenges introduced by tool invocation: tool-selection uncertainty, where an agent may invoke a competing tool instead of the target tool, and tool-argument compression, where fine-grained query information may be lost when the agent generates tool arguments. To tackle these challenges, we propose ToolSiphon, a query-only extraction attack that introduces two complementary signals: a target-discriminative signal, implemented through Tool Contrastive Analysis, to steer queries toward the target tool; and a response-grounded factual signal, implemented through Evidence Chained Feedback, to mitigate argument compression and progressively expand extraction coverage. Across three types of knowledge-based tools and six domain-specific datasets, ToolSiphon recovers 74.3% of source records on average when coarse-grained information about non-target tools is available, with 83.2% textual recovery and 90.2% semantic similarity. Even without such information, it recovers 66.3% of source records. ToolSiphon also remains effective against representative defenses and on three real-world agent platforms.