arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00839cs.CR

针对LLM提取的防御在不同攻击下是否有效?黑盒模型提取的生命周期基准

Do Defenses Against LLM Extraction Work Across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction

Shuze Liu, Kaixiang Zhao, Runyang Xu, Jingzhi Chen, Nathan Wu, Yu Wang, Yushun Dong

首次发表
浏览论文内容

中文总结 AI 辅助

提出统一基准,覆盖六种攻击、十种防御及两种自适应攻击,控制配置以评估黑盒模型提取防御的有效性。

中文摘要 AI 辅助

通过文本API部署的大型语言模型(LLMs)面临模型提取风险,因为攻击者可以收集其响应来训练能够复现其能力的替代模型。尽管已有研究开发了多种攻击和防御方法,但评估在访问假设、模型配置、查询预算和安全目标等方面仍然分散,限制了方法间的可比性。为解决这一问题,我们引入了一个统一基准,涵盖六种提取攻击、十种防御方法以及两种自适应攻击(在替代模型训练前对受保护的响应进行改写或回译)。该基准在每次比较中控制模型配置、查询数据、预算和保留评估条件,同时保留攻击特定的查询和训练流程。我们测量替代模型的能力、对受害模型的保真度、使用Rep-4的输出质量以及查询预算敏感性;防御方法使用自身的安全指标与替代模型性能相结合进行评估。对于自适应攻击,我们联合测量来源检测器得分以及基于改写响应训练的替代模型的能力和保真度。该基准为在纯文本访问条件下比较提取方法及其与防御的交互提供了可复现的基础。代码和工件可在该HTTPS URL获取。

英文摘要

Large language models (LLMs) deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. While prior work has developed diverse attacks and defenses, evaluations remain fragmented across access assumptions, model configurations, query budgets, and security objectives, limiting comparability across methods. To address this gap, we introduce a unified benchmark covering six extraction attacks, ten defenses, and two adaptive attacks that paraphrase or back-translate protected responses before surrogate training. The benchmark controls model configurations, query data, budgets, and held-out evaluation conditions within each comparison while preserving attack-specific querying and training procedures. We measure surrogate capability, fidelity to the victim, output quality using Rep-4, and query-budget sensitivity; defenses use their own security metrics paired with surrogate performance. For the adaptive attacks, we jointly measure provenance-detector scores and the capability and fidelity of surrogates trained on rewritten responses. The benchmark thus provides a reproducible basis for comparing extraction methods and their interactions with defenses under text-only access. Code and artifacts are available at https://github.com/sliu11-byte/MEA-Bench.

发表机构

  • Florida State University(佛罗里达州立大学)
  • Brigham Young University(杨百翰大学)
  • University of Michigan, Ann Arbor(密歇根大学安娜堡分校)
  • State University of New York at Buffalo(纽约州立大学布法罗分校)
  • Wake Forest University(维克森林大学)
  • University of Georgia(佐治亚大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑