多模态大语言模型指纹识别
Fingerprinting Multimodal Large Language Models
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多模态大语言模型易受非法部署和蒸馏侵权的问题,提出AttnPrint(基于跨模态注意力低频分量的白盒指纹)和DistillTrace(基于输出假设检验的黑盒审计),在154个模型实例上实现强检测性能与鲁棒性。
AI中文摘要:
尽管多模态大语言模型(MLLMs)支持广泛的图像-文本推理任务,但近期事件表明它们容易遭受非法部署和未经授权的蒸馏。现有的模型溯源解决方案通常因MLLMs中共享的语言骨干网络而受到干扰,并且难以检测蒸馏违规行为。为弥合这一差距并保护模型所有权,我们提出了关于多模态模型指纹识别的首项研究。受近期关于自注意力机制充当低通滤波器且其低频分量具有信息性的发现的启发,我们开发了AttnPrint用于白盒溯源。具体而言,我们提取跨模态注意力分布并分离其低频分量作为模型指纹。为促进黑盒审计,我们进一步引入了DistillTrace,该方法通过对MLLM输出进行假设检验来识别潜在的模型侵权。我们在涵盖19种多模态架构的154个模型实例上进行了广泛实验。值得注意的是,AttnPrint在实现强衍生模型检测性能的同时,对五种下游修改技术保持鲁棒性。DistillTrace还在三种参数无关技术下提供了蒸馏关系的证据。
英文摘要:
While multimodal large language models (MLLMs) enable a wide range of image-text reasoning tasks, recent incidents indicate that they are vulnerable to illicit deployment and unauthorized distillation. Existing solutions for model provenance are typically confounded by shared language backbones in MLLMs and struggle to detect violations of distillation. To bridge this gap and safeguard model ownership, we present the first study on multimodal model fingerprinting. Inspired by recent findings that self-attention acts as a low-pass filter and that its low-frequency components are informative, we develop AttnPrint for white-box provenance. Specifically, we extract cross-modal attention distributions and isolate their low-frequency components to serve as model fingerprints. To facilitate black-box auditing, we further introduce DistillTrace, which employs hypothesis testing of MLLM outputs to identify potential model infringement. We conduct extensive experiments on 154 model instances across 19 multimodal architectures. Notably, AttnPrint achieves strong derivative-model detection performance while remaining robust to five downstream modification techniques. DistillTrace also provides evidence of distillation relationships under three parameter-independent techniques.