CounterPersona:针对未经授权的人设技能蒸馏的仅追加防御
CounterPersona: Append-Only Defense Against Unauthorized Persona Skill Distillation
浏览论文内容
中文总结 AI 辅助
针对未经授权的人设技能蒸馏,提出仅追加防御方法CounterPersona,通过构建反人设证据和一致性重写,在多种度量下有效保护隐私与劳动自主权。
中文摘要 AI 辅助
人设技能蒸馏可以从个人信息中提取重复出现的模式,并将其编码为可复用的技能,从而使人工智能系统能够紧密复制个体的行为。然而,这种复制也引发了关于个人隐私和劳动自主权的严重担忧。与现有的基于扰动的防御方法不同,这些方法要求个体在数据收集之前修改其数据,而一旦历史记录被攻击者收集,就无法再被更改、清理或撤销。因此,此类防御难以适应这种仅追加的设置。为了解决这一挑战,我们引入了CounterPersona,它构建有针对性的反人设证据,将兼容的行为状态打包为紧凑的实现单元,并通过基于理由的一致性重写来强化它们。我们进行了大量实验,表明CounterPersona在词汇、语义和基于LLM的度量上均取得了强大且一致的效果,同时在不同蒸馏器上保持鲁棒性。我们的工作建立了一种技能反蒸馏范式,以保护个人隐私和劳动自主权免受未经授权的技能蒸馏侵害。
英文摘要
Persona skill distillation can extract recurring patterns from personal information and encode them into reusable skills, enabling AI systems to closely replicate an individual's behavior. However, such replication also raises serious concerns regarding personal privacy and labor autonomy. Unlike existing perturbation-based defenses that require individuals to modify their data before collection, once historical records are collected by an attacker, they can no longer be altered, sanitized, or revoked. Therefore, such defenses are difficult to adapt to this append-only setting. To solve this challenge, we introduce CounterPersona, which constructs targeted counter-persona evidence, packs compatible behavioral states into compact realization units, and strengthens them through rationale-guided consistency rewriting. We conduct extensive experiments showing that CounterPersona achieves strong and consistent effectiveness across lexical, semantic, and LLM-based measures, while remaining robust across distillers. Our work establishes a skill anti-distillation paradigm for protecting personal privacy and labor autonomy against unauthorized skill distillation.
发表机构
- University of Electronic Science and Technology of China(电子科技大学)
- The University of Hong Kong(香港大学)
- City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。