发表机构
Khoury College of Computer Sciences; Northeastern University(胡里计算机学院; 东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM工具调用带来的隐私风险,提出AgentDOXX评估套件,通过822份合成访谈记录量化重识别攻击,发现检索与参数记忆共同作用,并利用攻击轨迹指导匿名化。
AI 中文摘要
随着大型语言模型(LLM)获得网络搜索等工具使用能力,它们能够检索并交叉引用公开信息,从而产生超越记忆效应的隐私风险。其一种表现形式是重识别:将匿名访谈记录与具名个体关联起来。然而,在没有真实身份标签的情况下,此类攻击的覆盖范围以及防御措施所提供的保护无法被可靠衡量。我们引入了AgentDOXX,一个包含822份合成访谈记录的评估套件,这些记录基于具有已知身份的真实个体的公开信息构建。我们评估了十五种开源权重模型和专有模型的配置,隔离了网络搜索的影响,并分析了它们的搜索轨迹以区分检索驱动型识别与参数记忆型识别。真实身份标签揭示,重识别风险分布于智能体的执行过程中:检索和参数记忆均有所贡献,其中开源权重模型在无搜索情况下识别了15-28%的记录;一旦目标出现在检索结果中,识别成功率超过88%;实体掩码在分层样本上仍使至少一名攻击者在85.3%的案例中成功;隐私指令抑制了命名但未抑制检索,配置在准确率为0%的情况下仍在多达62%的记录中检索到了目标主体。我们进一步表明,观察到的攻击轨迹可为定位识别性片段提供监督,从而为攻击感知的匿名化提供一条路径。
英文摘要
As Large Language Models (LLMs) gain tool use capabilities such as web search, they can retrieve and cross-reference public information, creating privacy risks beyond memorization. One manifestation is re-identification: linking an anonymized interview transcript to a named individual. Yet without ground-truth identities, the coverage of such attacks and the protection offered by a defense cannot be reliably measured. We introduce AgentDOXX, an evaluation suite of 822 synthetic interview transcripts grounded in public information about real individuals with known identities. We evaluate fifteen configurations of open-weight and proprietary models, isolating the effect of web search, and analyze their search trajectories to distinguish retrieval-driven from parametric identifications. Ground-truth identities reveal that re-identification risk is distributed across an agent's execution: retrieval and parametric recall both contribute, with open-weight models identifying 15-28% of transcripts without search; identification succeeds in over 88% of cases once the target appears in a retrieved result; entity masking leaves at least one attacker successful on 85.3% of a stratified sample; and privacy instructions suppress naming but not retrieval, with configurations scoring 0% accuracy yet retrieving the subject in up to 62% of transcripts. We further show that observed attack trajectories can provide supervision for localizing identifying spans, offering a path toward attack-informed anonymization.