从个性化大语言模型(LLM)智能体的行为中推断隐藏的用户模型
Inferring Hidden User Models from the Behavior of Personalized LLM Agents
- The Hong Kong Polytechnic University(香港理工大学)
- Huawei(华为)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对个性化LLM智能体的用户模型提出黑盒攻击方法UMPeek,经基准评估和现实系统验证,其优于现有攻击,可在响应级防御下恢复用户信息,揭示了相关隐私风险。
AI中文摘要:
近年来,个性化大语言模型(LLM)智能体越来越多地将记忆中保留的信息转化为压缩或结构化的表示,我们将其称为用户模型,以指导后续决策。当普通接口可访问的状态中移除源文本时,这些模型通常被视为更具隐私保护性,因为直接的内存提取攻击会失去其目标文本。然而我们认为,用户模型暴露了新的攻击面,因为即使源记录和后端状态仍无法访问,攻击者仍可从用户模型所塑造的个性化选择中恢复私人信息。因此,我们提出UMPeek,这是一种基于假设引导的自适应探测的黑盒攻击,用于推断此类隐藏的用户模型。它从请求留下的开放选择中形成假设,在普通的后续任务间切换,仅保留可见行为支持且不矛盾的主张。我们针对不同的个性化任务和用户模型后端,对现有攻击进行了广泛的基准评估,还使用已确认保留的信息在现实系统中验证了UMPeek,并评估了针对其自适应探测的防御措施。总体而言,在基准和现实世界比较中,UMPeek的表现优于现有攻击,且在响应级防御下仍能恢复用户信息,这表明当保留的信息塑造可见行为时,使记录和后端状态无法访问并不能保证语义隐私。
英文摘要:
Recent personalized LLM agents increasingly transform information retained in memory into compressed or structured representations, which we call user models, to guide later decisions. When source wording is removed from the state reachable through the ordinary interface, these models are commonly treated as more privacy-preserving because direct memory-extraction attacks lose the text they target. Yet we argue that user models expose a new attack surface because an attacker can still recover the private information from the personalized choices they shape, even when source records and backend state remain inaccessible. We therefore introduce UMPeek, a black-box attack based on hypothesis-guided adaptive probing to infer such hidden user model. It forms hypotheses from choices left open by a request, switches among ordinary follow-up tasks, and retains only claims supported and not contradicted by visible behavior. We conduct an extensive benchmark evaluation across diverse personalization tasks and user-model backends against existing attacks. We further validate UMPeek in real-world systems using information confirmed to be retained, and we evaluate defenses against its adaptive probing. Overall, UMPeek outperforms existing attacks in both benchmark and real-world comparisons and continues to recover user information under response-level defenses, showing that keeping records and backend state inaccessible does not guarantee semantic privacy when retained information shapes visible behavior.