arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ContextLeak:通过恶意工具窃取大语言模型智能体的上下文

ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools

Yuqi Jia, Ruiqi Wang, Patrick Li, Yuepeng Hu, Peinian Li, Neil Gong

arXiv 2608.27800首次发表:更新:

AI 中文总结

本研究提出ContextLeak恶意工具攻击,通过强化学习设计工具名称与描述诱导LLM智能体泄露上下文,其在影子用户与受害者上下文差异大时仍保持高有效性,且性能优于现有同类攻击。

AI 中文摘要

窃取大语言模型(LLM)智能体的运行时上下文,例如用户提示、执行轨迹和工具列表,会对用户造成严重的安全和隐私风险。此类攻击可通过恶意工具实施,通常需要满足三个条件:(1)智能体选择该恶意工具执行任务;(2)智能体将其运行时上下文作为输入参数传递给该工具;(3)工具的实现代码将这些输入传输到攻击者控制的端点。现有研究主要关注条件(1)和(3),却在很大程度上忽略了对成功实施上下文窃取至关重要的条件(2)。在本研究中,我们通过开发ContextLeak来填补这一空白,这是一种恶意工具攻击,可诱导智能体同时选择该工具并将其上下文作为输入参数泄露。我们通过使用强化学习精心设计工具的名称和描述来实现这一攻击。具体而言,ContextLeak使用一个名为攻击LLM的大语言模型来自动生成恶意工具的名称和描述。为提高攻击有效性,我们在一组具有多样化、模拟智能体上下文的影子用户上,通过强化学习对攻击LLM进行微调。我们的关键技术贡献是设计了针对上下文窃取目标量身定制的新型奖励函数,使基于强化学习的攻击LLM微调能够有效进行。广泛的评估表明,即使影子用户的上下文与受害者用户的上下文存在显著差异,我们的攻击仍然保持高度有效性。此外,在适配该场景后,ContextLeak的性能显著优于现有的恶意工具攻击。

英文摘要

Exfiltrating an LLM agent's runtime context -- such as the user prompt, execution trajectory, and tool list -- poses severe security and privacy risks to users. Such attacks can be carried out via malicious tools and typically require three conditions: (1) the agent selects the malicious tool for task execution, (2) the agent passes its runtime context as input arguments to the tool, and (3) the tool's implementation transmits these inputs to an attacker-controlled endpoint. Existing work primarily focuses on conditions (1) and (3), leaving condition (2) largely unexplored, despite its critical role in enabling successful context exfiltration. In this work, we bridge this gap by developing ContextLeak, a malicious tool attack that induces the agent to both select the tool and disclose its context as input arguments. We realize this attack by carefully crafting the tool's name and description using reinforcement learning. Specifically, ContextLeak employs an LLM, referred to as the attack LLM, to automatically generate the malicious tool's name and description. To improve attack effectiveness, we fine-tune the attack LLM via reinforcement learning on a set of shadow users with diverse, simulated agent contexts. Our key technical contribution is the design of novel reward functions tailored to the context exfiltration objective, enabling effective reinforcement-learning-based fine-tuning of the attack LLM. Extensive evaluation demonstrates that our attack remains highly effective even when the shadow users' contexts differ substantially from those of the victim users. Moreover, ContextLeak significantly outperforms existing malicious tool attacks when adapted to this setting.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑