发表机构
Minjiang University; Renmin University of China; RMIT University(闽江学院; 中国人民大学; 皇家墨尔本理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CIPL提出一种通道感知评估框架,通过多阶段建模统一比较LLM智能体中隐私泄露的可恢复性,实验揭示存储标签不决定泄露,为异构流水线提供通用比较基准。
AI 中文摘要
LLM智能体中的隐私泄露通常仅在单个组件(如记忆、检索或工具使用流水线)内进行评估,这使得难以区分内部暴露与外部观察者实际可恢复的信息。我们提出CIPL(通道反演用于隐私泄露),一种用于LLM智能体黑盒隐私泄露的通道感知评估框架。CIPL通过敏感源、选择、组装、执行、观察和提取阶段来表示目标,并在共享协议下评估从选定敏感单元到攻击者可恢复输出的转变。针对基于记忆、检索中介和工具中介的目标进行的实验,以及一个BrowserUse实时智能体案例研究,表明存储标签本身并不决定可恢复性。记忆目标构成近乎饱和的参考案例,检索中介的泄露通常是部分的,而工具中介和实时智能体的泄露随观察表面、提示到通道的对齐、检索深度和提供商行为而强烈变化。分层语义审计进一步识别了规范精确匹配所遗漏的攻击者有用的披露。因此,CIPL提供了一个通用框架,用于比较异构智能体流水线中内部敏感依赖如何实现为外部可恢复泄露。
英文摘要
Privacy leakage in LLM agents is commonly evaluated within individual components such as memory, retrieval, or tool-use pipelines, which makes it difficult to distinguish internal exposure from information that an external observer can actually recover. We present CIPL (Channel Inversion for Privacy Leakage), a channel-aware evaluation framework for black-box privacy leakage in LLM agents. CIPL represents a target through sensitive source, selection, assembly, execution, observation, and extraction stages and evaluates the transition from selected sensitive units to attacker-recoverable output under a shared protocol. Experiments across memory-based, retrieval-mediated, and tool-mediated targets, together with a BrowserUse live-agent case study, show that storage labels alone do not determine recoverability. Memory targets form a near-saturated reference case, retrieval-mediated leakage is frequently partial, and tool-mediated and live-agent leakage varies strongly with observation surface, prompt-to-channel alignment, retrieval depth, and provider behavior. A stratified semantic audit further identifies attacker-useful disclosures that canonical exact matching misses. CIPL therefore provides a common framework for comparing how internal sensitive dependence is realized as externally recoverable leakage across heterogeneous agent pipelines.
Comments58 pages, 4 figures; includes appendix