AI 中文总结
研究模型窃取攻击,提出受水印技术启发的防御方法,通过扰动对数its层防止攻击,经实证实验验证防御有效性,在保留模型效用的同时有效抵御攻击。
AI 中文摘要
模型窃取攻击最近出现了,能从黑盒商业语言模型中提取精确信息。本文提出针对\cite{carlini2024stealing}近期攻击的防御方法及用于提取生产语言模型隐藏层维度的扩展方法。方法受水印技术启发,通过扰动模型的对数its层来防止攻击。提供实证实验证明防御有效性及对模型质量退化的影响,提出有效防御且保留模型效用。
英文摘要
Model stealing attacks have recently been introduced, enabling the extraction of precise information from black-box commercial language models. In this work, we propose defense methods against a recent attack of \cite{carlini2024stealing} and extensions for extracting the hidden layer dimension of production language models. Our methods are inspired by watermarking techniques that perturb the logits layer of these models to prevent such attacks. We provide empirical experiments demonstrating the effectiveness of the proposed defense versus model quality degradation across various configurations, and propose an effective defense against such attacks while preserving model utility.