显而易见的爪痕:通过大语言模型智能体工具调用实现未授权的上下文信息泄露
The Claws in Plain Sight: Unauthorized Context Disclosure through LLM Agent Tool Calls
浏览论文内容
中文总结 AI 辅助
该研究提出Claw in Plain Sight攻击方法,通过构造任务相关内容框架诱导LLM智能体泄露上下文信息,实验显示其在多种模型上的会话级泄露率达20.8%-75.0%,提示需在工具参数执行前进行目的与目的地感知检查。
中文摘要 AI 辅助
大语言模型(LLM)智能体通常会基于用户配置文件、对话历史、检索到的文档以及先前的工具结果来构建工具调用参数。然而,合法访问上下文信息并不意味着可以出于任何目的或向任何目的地传输该信息。我们提出了Claw in Plain Sight(显而易见的爪痕),这是一种权威压力攻击,其中与任务相关的内容框架将受保护属性表述为操作或程序上的必要要求,从而导致模型将这些属性纳入原本有效的生成参数中。我们使用受控的合成基准对Claw in Plain Sight进行评估,该基准在五种DeepSeek和Claude模型配置中,将六个压力级别与四个隐私策略级别交叉组合,生成了120次调用。在完整的压力-策略矩阵中,测试模型的会话级泄露率范围为20.8%至75.0%。更严格的隐私指令会降低总体泄露率,但并不能在所有模型中一致地消除泄露,这表明提示级策略无法提供可移植的执行边界。我们的实验仅使用合成配置文件并在本地捕获生成的参数;它们测量的是上下文到参数边界处违反策略的生成,而非已完成的网络数据 exfiltration(泄露)或部署用户的信息泄露。这些发现促使在工具调用执行前,对生成的工具参数进行感知目的和目的地的检查。
英文摘要
LLM agents routinely construct tool-call arguments from user profiles, conversation history, retrieved documents, and prior tool results. However, legitimate access to contextual information does not imply authorization to transmit that information for every purpose or destination. We present Claw in Plain Sight, an authority- pressure attack in which task-adjacent content frames protected attributes as operationally or procedurally required, causing a model to include them in otherwise valid generated arguments. We evaluate Claw in Plain Sight using a controlled synthetic benchmark that crosses six pressure levels with four privacy-policy levels across five DeepSeek and Claude model configurations, producing 120 calls. Across the complete pressure-policy matrix, session-level disclosure rates range from 20.8% to 75.0% among the tested models. Stronger privacy instructions reduce aggregate disclosure but do not eliminate it consistently across models, showing that prompt-level policies do not provide a portable enforcement boundary. Our experiments use only synthetic profiles and capture proposed arguments locally; they measure policy-violating generation at the context-to-argument boundary, not completed network exfiltration or leakage from deployed users. These findings motivate purpose- and destination-aware inspection of generated tool arguments before execution.