发表机构
University of California, Irvine; University of California, Santa Cruz; University of California, Davis; University of California, Santa Barbara(加州大学欧文分校; 加州大学圣克鲁兹分校; 加州大学戴维斯分校; 加州大学圣巴巴拉分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体隐私泄露,提出三种无需提示注入的攻击,并设计工具级拦截器FLOWSEAL,基于数据来源与信息流控制格,将泄露率降至近零且不损任务效用。
AI 中文摘要
基于大型语言模型(LLM)构建的个人AI智能体,为了提供个性化协助,越来越多地被授予访问用户私人数据和通信的权限。这种访问权限带来了持续性的隐私风险:智能体必须决定某一特定敏感信息是否应向特定方披露。现有防御措施通过更强大的系统提示、训练或显式同意检查程序,使智能体的后端LLM更具隐私保护性,但这种方法存在结构性挑战:每当强制执行是LLM在攻击者控制的同一对话上下文中做出的判断时,强制执行机制与攻击面便重合了。我们通过三种仅需普通智能体交互、无需提示注入的新攻击,证明了现有防御的不足:协作工作区诱饵(Collaborative Workspace Lure)将提取尝试重新框定为协作工作;语义混淆攻击(Semantic Obfuscation Attack)通过省略而非智能体所写内容诱导披露;通道解耦攻击(Channel Decoupling Attack)将提取请求与披露跨独立通道拆分。这三种攻击的泄露率均显著高于这些防御最初设计所能承受的攻击。基于此观察,我们提出了FLOWSEAL,一种防御机制,通过在LLM上下文之外的工具级拦截器强制执行机密性,其基础是数据来源和带有受控降密的信息流控制格。在三个基准、五个基于提示的基线和八种攻击(包括通过MCP执行实时工具调用的真实智能体)上的评估表明,FLOWSEAL将泄露率降至接近零(例如,针对协作工作区诱饵,从52.2%降至0.5%),同时保持任务效用,且与底层LLM后端无关。
英文摘要
Personal AI agents built on large language models (LLMs) are increasingly given access to a user's private data and communications in order to provide personalized assistance. This access creates a persistent privacy risk: the agent must decide whether a given sensitive information should be disclosed to a particular party. Existing defenses address this by making the agent's backend LLM more privacy-preserving through stronger system prompts, training, or explicit consent-checking procedures, but this approach has a structural challenge: whenever enforcement is a judgment the LLM makes over the same conversational context an adversary controls, the enforcement mechanism and the attack surface coincide. We demonstrate this against existing defenses with three new attacks that require only ordinary agent interaction and no prompt injection: Collaborative Workspace Lure reframes an extraction attempt as collaborative work; Semantic Obfuscation Attack induces disclosure through omission rather than through anything the agent writes; and Channel Decoupling Attack splits the extraction request and the disclosure across independent channels. All three achieve substantially higher leak rates than the attacks these defenses were originally designed to withstand. Guided by this observation, we present FLOWSEAL, a defense that enforces confidentiality through a tool-level interceptor outside the LLM's context, grounded in data provenance and an information-flow-control lattice with controlled declassification. Evaluated across three benchmarks, five prompt-based baselines, and eight attacks, including a real agent executing live tool calls through MCP, FLOWSEAL reduces leak rates to near zero (e.g., 52.2% to 0.5% against Collaborative Workspace Lure) while preserving task utility, regardless of the underlying LLM backend.