当智能体学会成为你:对角色技能中的隐私泄露、冒充风险及防御措施进行基准测试
When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
浏览论文内容
中文总结 AI 辅助
该研究推出AntiSkillBench基准,评估角色技能流程的隐私泄露、冒充风险及防御,发现风险持续存在,现有防御效果有限,为相关技能开发提供基准。
中文摘要 AI 辅助
角色技能将个人互动历史提炼为可移植、可执行的产物,供下游智能体使用。在实现灵活个性化的同时,这一过程会集中碎片化的个人信号,通过重复使用放大其影响,并对针对单个记录或基于检索的记忆设计的防御措施构成挑战。为系统研究角色技能流程的安全性,我们推出AntiSkillBench,这是一个用于评估角色技能流程全链路风险与防御措施的端到端基准。它包含三部分:(i)7500条基于角色的对话轨迹数据集,由涵盖不同任务场景的50个行为丰富的个人资料构建而成;(ii)评估套件,用于在三种技能提炼策略下测量技能级隐私泄露、智能体级属性泄露及行为冒充情况;(iii)防御评估,涵盖在线与事后干预的四种配置,包括主动风险抑制与被动来源保护。对三个前沿智能体开展的实验表明,角色技能风险在智能体主干与提炼协议中持续存在,从显式属性延伸至沟通风格与人格特质。现有防御措施的效果有限且依赖提炼方式,无法在风险与提炼策略间实现泛化。这些结果凸显AntiSkillBench是开发隐私保护与真实性感知角色技能的具有挑战性的基准。
英文摘要
Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges defenses designed for individual records or retrieval-based memory. To systematically investigate the safety of the persona-skill pipeline, we introduce AntiSkillBench, an end-to-end benchmark for evaluating risks and defenses across the persona-skill pipeline. It comprises: (i) a dataset of 7,500 persona-grounded dialogue traces, constructed from 50 behaviorally rich profiles spanning diverse task scenarios; (ii) an evaluation suite that measures skill-level privacy leakage and agent-level attribute disclosure and behavioral impersonation across three skill-distillation strategies; and (iii) a defense evaluation covering four configurations across online and post-hoc interventions, including active risk suppression and passive provenance protection. Experiments across three frontier agents show that persona-skill risks persist across agent backbones and distillation protocols, extending from explicit attributes to communication styles and personality traits. Existing defenses exhibit limited and distillation-dependent effectiveness, failing to generalize across risk and distillation strategies. These results highlight AntiSkillBench as a challenging benchmark for developing privacy-preserving and authenticity-aware persona skills.
发表机构
- The University of Sydney(悉尼大学)
- University of Queensland(昆士兰大学)
- Southeast University(东南大学)
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。