arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.18829cs.LGcs.CR

无损反蒸馏采样

Lossless Anti-Distillation Sampling

  • Peking University(北京大学)
  • Tsinghua University(清华大学)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Zibo Diao, Jingchu Gai, Xinyue Ai, Zhang Zhang, Zhenyu He, Di He

AI总结:

本文提出了一种无损反蒸馏采样方法,通过在保持良性用户体验的同时,有效对抗多账号蒸馏攻击,降低蒸馏模型的泛化能力。

AI中文摘要:

面向商业生成模型的前沿领域,蒸馏攻击正成为日益严峻的威胁。蒸馏者通过收集生成响应并以极低的成本训练自己的竞争模型。现有防御措施要么依赖于修改模型输出,从而牺牲良性用户的响应质量,要么依赖于行为检测方法,这些方法可以通过在多个账户上分布查询来轻易绕过。在本工作中,我们提出了无损反蒸馏采样(LADS),一种专门设计用于对抗多账号蒸馏同时保持良性用户体验的新型采样方案。具体而言,LADS从由查询的语义内容和用户查询次数决定的私有种子中推导出每种生成的随机性。通过构造,每个良性用户在每次访问时都会独立地从原始模型中采样响应,因此不会产生失真。相反,对于蒸馏者,不同账户在相同语义桶中的查询会共享潜在随机性。因此,收集的数据变得相关,可能降低样本多样性并损害泛化能力。利用统一收敛理论,我们证明LADS在无条件和条件生成设置中,能够证明降低蒸馏者泛化差距的收敛率相对于标准i.i.d.采样。在图像生成、数学推理和代码生成的实验中,证实LADS显著降低蒸馏学生的表现,同时保持对单个用户的精确统计保真度。

英文摘要:

Frontier commercial generative models face a growing threat from distillation, whereby a distiller harvests generated responses and trains a competing model at drastically lower cost. Existing defenses either modify the generation to degrade distillation performance, sacrificing response quality, or rely on behavioral detection mechanisms that can be readily bypassed through multi-account querying. In this work, we propose Lossless Anti-Distillation Sampling (LADS), which leaves the generation itself unchanged while substantially reducing the effectiveness of distillation. Concretely, LADS controls the latent randomness underlying inference through a coupling mechanism that preserves within-account generation independence while inducing cross-account dependence. By construction, each benign user, who typically holds only a single account, receives the same experience under LADS as they would without any defense, thereby enjoying a lossless experience. However, for a multi-account task-specific distiller, semantically similar queries submitted across different accounts are assigned coupled randomness, inducing dependence in the harvested data and thereby degrading the generalization performance of the distilled model. Using uniform convergence theory, we show that LADS provably degrades the distiller's generalization gap relative to standard i.i.d. sampling. Experiments on image generation, mathematical reasoning, and code generation confirm that LADS substantially degrades the performance of distilled students while preserving exact statistical fidelity for individual users.

↑