面向黑盒大语言模型所有权验证的针对性反事实指纹技术
Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification
查看机构详情
- Huazhong University of Science and Technology(华中科技大学)
- Alibaba Group(阿里巴巴集团)
- Microsoft Corporation(微软公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对黑盒LLM所有权验证的挑战,提出TCF框架,将开放式生成对比转化为受限答案的反事实迁移,在四个LLM系列上平均AUC达0.9861,优于TRAP等现有方法。
中文摘要 AI 辅助
大语言模型(LLMs)是高价值资产,可通过重新部署、微调、量化或进一步对齐得到。由于部署后的LLMs通常仅通过查询API暴露,所有权验证往往依赖黑盒文本响应。该场景存在挑战:生成内容是开放式的,且重复查询结果可能存在差异,而现有黑盒指纹依赖的信号在最终响应接口下脆弱,包括全文匹配、软行为特征或设计为不可迁移的模型特定提示。我们提出TCF(Targeted Counterfactual Fingerprinting,针对性反事实指纹技术),这是一种黑盒LLM指纹框架,将开放式生成对比转化为受限答案的针对性反事实迁移。TCF将每个验证查询限制在有限答案空间内,减少进入验证分数的表面形式歧义,并优化提示扰动,使其朝向与受保护模型在原始提示下的干净答案不同的反事实目标。验证过程简化为检查嫌疑模型解析后的最终答案是否与记录的目标匹配。我们引入源模型反事实边际(SCM),这是仅受保护模型具备的量,用于在扰动前验证目标是否不可能,扰动后是否可能;SCM控制目标选择、扰动停止和指纹过滤。在由局部行为接近性驱动的显式派生保留和独立迁移预算下,我们推导了派生模型与独立模型之间的目标准确率差距。在四个LLM系列上,TCF实现了平均AUC为0.9861,比TRAP、ProFLingo和ZeroPrint提升了0.07至0.19。
英文摘要
Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment. Because deployed LLMs are commonly exposed only through query APIs, ownership verification must often rely on black-box text responses. This setting is difficult: generations are open-ended and can vary across repeated queries, while existing black-box fingerprints rely on signals that are fragile under a final-response interface, including full-text matching, soft behavioral features, or model-specific prompts designed not to transfer. We propose TCF (Targeted Counterfactual Fingerprinting), a black-box LLM fingerprinting framework that converts open-ended generation comparison into constrained-answer targeted counterfactual transfer. TCF restricts each verification query to a finite answer space, reducing the surface-form ambiguity that enters the verification score, and optimizes a prompt perturbation toward a counterfactual target different from the protected model's clean answer on the original prompt. Verification reduces to checking whether the suspect model's parsed final answer matches the recorded target. We introduce the source-model counterfactual margin (SCM), a protected-model-only quantity that certifies the target is unlikely before the perturbation and likely after it; SCM controls target selection, perturbation stopping, and fingerprint filtering. Under explicit derived-preservation and independent-transfer budgets motivated by local behavioral closeness, we derive a target-accuracy gap between derived and independent models. Across four LLM families, TCF achieves an average AUC of 0.9861, improving over TRAP, ProFLingo, and ZeroPrint by 0.07 to 0.19.