发表机构
Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对基础模型版权侵权检测问题,提出统一的DCS框架,将侵权证据视为反事实条件分布转移,通过条件差分隐私形式化,创建双分支并结合多种因素界定敏感度,还定义校准统计量,适用于多种模型并以不同方式评估。
AI 中文摘要
当前,多数基础模型能重现或强烈依赖受版权保护的训练内容,但仅输出相似度不足以进行侵权检测,因为相似输出可能源于公共领域概念等。本文开发了一个统一的事后检测框架,将版权侵权证据视为反事实条件分布转移。通过条件差分隐私形式化此观点并引入双分支条件敏感度(DCS)。具体而言,DCS框架围绕已部署模型创建学习和反学习分支,通过影响函数分析连接其位移与不可得的反事实再训练效果,并用反事实隐私预算替代物等界定可观测敏感度。还定义了校准检测统计量以区分特定目标记忆与一般微调不稳定性。该框架适用于多种模型,通过不同方式进行评估。
英文摘要
Currently, most foundation models can reproduce or strongly depend on copyrighted training content, but output similarity alone is insufficient for infringement detection, because similar outputs may also arise from public-domain concepts, common stylistic conventions, or ordinary statistical generalization. In this paper, we develops a unified post-hoc detection framework that treats copyright infringement evidence as a counterfactual conditional distribution shift: a protected target is suspicious when the model's behavior under aligned conditions would change measurably if that target were included in, or removed from, the training process. We formalize this view through conditional differential privacy and introduce Dual-Branch Conditional Sensitivity (DCS), an operational statistic that measures the observable gap between two locally perturbed model states. Specifically, the proposed DCS framework creates a learning branch and an unlearning branch around the deployed model, connects their displacement to the unavailable counterfactual retraining effect through influence-function analysis, and bounds the observable sensitivity by the counterfactual privacy-budget surrogate, local curvature, training-set scale, and perturbation step size. To distinguish target-specific memorization from generic fine-tuning instability, we further define a calibrated detection statistic that subtracts the sensitivity measured under orthogonal conditions. The DCS framework is instantiated for ridge-regularized linear regression, conditional diffusion models, autoregressive language models, and multimodal models. These instantiations show how the same principle can be evaluated through prediction gaps, image-embedding divergence, token-distribution or entropy shifts, and cross-modal representation changes.