CopyShield:大语言模型中版权防御的跨级别基准
CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs
浏览论文内容
中文总结 AI 辅助
本研究提出跨级别基准CopyShield,对比三种不同干预级别的大语言模型版权防御方法,在两类模型上揭示其合规性-实用性权衡,发现针对性非逐字抑制为待解决挑战。
中文摘要 AI 辅助
大语言模型可以逐字重现记忆的文本,但版权防御通常在不兼容的协议下进行评估。我们推出CopyShield,这是一个受控基准,用于比较三种不同干预级别的代表性防御方法:对比解码(输出层面)、直接偏好优化(DPO,行为层面)和激活干预(表示层面)。我们在两个模型家族LLaMA-3.1-8B和Mistral-7B-v0.3上评估CopyShield,使用对五本公共领域书籍的受控记忆,并采用共享协议测量逐字泄漏、校准后的非逐字泄漏、实用性和退化性。在这些方法中,干预级别与不同的合规性-实用性权衡相关。在LLaMA-3.1-8B上,对比解码几乎无退化(0-2%),但在NV-Recall为0.192-0.203时达到逐字抑制下限;DPO几乎消除了逐字泄漏(从0.263降至0.002),但在58%的问答输出中引发 paraphrase-loop 退化,且相较于SFT基线无实用性提升;激活干预通过在生成前阻止84%的非逐字查询,实现了最低的非逐字标记率(1/200)。人工评估证实,DPO的连贯性较低,而激活干预通过广泛的弃权(不执行)降低了感知到的版权风险。在Mistral-7B-v0.3上,输出和表示层面的模式持续存在,而DPO退化降至10-14%,表明其严重程度依赖于模型。综上,CopyShield提供了跨级别参考基准,并确定针对性的非逐字抑制是一个未解决的挑战。代码可在此URL获取。
英文摘要
Large language models can reproduce memorized text verbatim, yet copyright defenses are usually evaluated under incompatible protocols. We introduce CopyShield, a controlled benchmark comparing three representative defenses at distinct intervention levels: contrastive decoding (output), Direct Preference Optimization (behavioral), and activation intervention (representation). We evaluate CopyShield on two model families, LLaMA-3.1-8B and Mistral-7B-v0.3, using controlled memorization over five public-domain books and a shared protocol measuring literal leakage, calibrated non-literal leakage, utility, and degeneracy. Across these methods, intervention level is associated with distinct compliance-utility trade-offs. On LLaMA-3.1-8B, contrastive decoding remains near-degeneracy-free (0-2%) but reaches a literal-suppression floor at NV-Recall 0.192-0.203. DPO nearly eliminates literal leakage (0.263 to 0.002) but induces paraphrase-loop degeneracy in 58% of QA outputs, with no utility gain over the SFT baseline. Activation intervention attains the lowest non-literal flagging rate (1/200) by blocking 84% of non-literal queries before generation. Human evaluation confirms that DPO has low coherence, whereas activation lowers perceived copyright risk through broad refusal. On Mistral-7B-v0.3, the output- and representation-level patterns persist, while DPO degeneracy falls to 10-14%, showing that its severity is model-dependent. Together, CopyShield provides cross-level reference baselines and identifies targeted non-literal suppression as an open challenge. The code is available at https://github.com/spotai-mbzuai/CopyShield.git.
发表机构
- Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学(MBZUAI))
机构由 AI 辅助整理,请以论文原文为准。