arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.22521cs.CRcs.AIcs.LG

大型语言模型模型提取攻击与防御综述

A Survey on Model Extraction Attacks and Defenses for Large Language Models

发表机构圣约翰大学 · 佛罗里达州立大学 · 西北大学
另 2 家 · 查看机构详情
  • University of Notre Dame(圣约翰大学)
  • Florida State University(佛罗里达州立大学)
  • Northwestern University(西北大学)
  • Duke University(杜克大学)
  • University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

Kaixiang Zhao, Lincan Li, Kaize Ding, Neil Zhenqiang Gong, Yue Zhao, Yushun Dong

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本综述系统梳理了针对大型语言模型的模型提取攻击与防御方法,提出分类体系与评估指标,并指出集成攻击与自适应防御等未来研究方向。

中文摘要 AI 辅助

模型提取攻击对已部署的语言模型构成重大安全威胁,可能损害知识产权和用户隐私。本综述提供了针对LLM的提取攻击与防御的全面分类体系,将攻击分为功能提取、训练数据提取和提示词定向攻击三类。我们分析了多种攻击方法,包括基于API的知识蒸馏、直接查询、参数恢复以及利用Transformer架构的提示词窃取技术。随后,我们考察了按模型保护、数据隐私保护和提示词定向策略组织的防御机制,并评估了它们在不同部署场景下的有效性。我们提出了用于评估攻击有效性和防御性能的专门指标,以应对生成式语言模型的特定挑战。通过分析,我们识别了当前方法中的关键局限性,并提出了有前景的研究方向,包括集成攻击方法和在安全性与模型效用之间取得平衡的自适应防御机制。本工作服务于NLP研究人员、机器学习工程师以及安全专业人士,旨在保护生产环境中的语言模型。

英文摘要

Model extraction attacks pose significant security threats to deployed language models, potentially compromising intellectual property and user privacy. This survey provides a comprehensive taxonomy of LLM-specific extraction attacks and defenses, categorizing attacks into functionality extraction, training data extraction, and prompt-targeted attacks. We analyze various attack methodologies including API-based knowledge distillation, direct querying, parameter recovery, and prompt stealing techniques that exploit transformer architectures. We then examine defense mechanisms organized into model protection, data privacy protection, and prompt-targeted strategies, evaluating their effectiveness across different deployment scenarios. We propose specialized metrics for evaluating both attack effectiveness and defense performance, addressing the specific challenges of generative language models. Through our analysis, we identify critical limitations in current approaches and propose promising research directions, including integrated attack methodologies and adaptive defense mechanisms that balance security with model utility. This work serves NLP researchers, ML engineers, and security professionals seeking to protect language models in production environments.

↑