系统提示词是否会留下行为指纹?一项基于输出相似度的克隆检测大规模实证研究
Do System Prompts Leave Behavioral Fingerprints? A Large-Scale Empirical Study of Clone Detection via Output Similarity
浏览论文内容
中文总结 AI 辅助
本研究提出黑盒行为指纹识别(BBF)方法,通过大规模实验验证其可检测商用LLM系统提示词的克隆,揭示了提示词选择对输出方差的影响及相关性能规律。
中文摘要 AI 辅助
商用大语言模型(LLM)的系统提示词可被以超过80%的成功率提取并零成本复用,但提示词所有者无法验证疑似部署是否为克隆。我们提出黑盒行为指纹识别(BBF):提示词所有者从模型输出中注册行为签名,后续测试疑似部署是否比无关基线更匹配该签名,BBF仅需黑盒API访问。通过大规模研究(4个模型家族、8个基准、288000条响应),我们发现提示词选择可解释24.4%的输出方差,同模型检测达到AUC 0.876;跨模型性能受检测器身份限制,非对角AUC范围为0.845(Claude作为检测器)至0.665(Qwen),整体均值0.725。BBF可抵御非适应性提示词改写(AUC≥0.889)且对不完善提取具有鲁棒性,但单句正式语气前缀会使短结构化输出检测崩溃(MNLI 0.978→0.547),凸显风格不变检测为关键开放问题;零成本查询选择规则诊断查询优化可使跨模型AUC提升0.120。
英文摘要
System prompts can be extracted from commercial LLMs with over 80\% success and redeployed at zero cost, yet a prompt owner has no way to verify whether a suspected deployment is a clone. We propose Black-Box Behavioral Fingerprinting (BBF): the prompt owner registers a behavioral signature from model outputs and later tests whether a suspect deployment matches that signature more closely than an unrelated baseline. BBF requires only black-box API access. Through a large-scale study (4 model families, 8 benchmarks, 288{,}000 responses), we find that prompt choice explains 24.4\% of output variance and same-model detection reaches AUC 0.876. Cross-model performance is bounded by detector identity, with off-diagonal AUC ranging from 0.845 (Claude as detector) down to 0.665 (Qwen) and overall mean 0.725. BBF resists non-adaptive prompt paraphrasing (AUC $\geq 0.889$) and is robust to imperfect extraction, but a single-sentence formal-tone prefix can collapse detection on short structured outputs (MNLI 0.978 $\to$ 0.547), isolating style-invariant detection as the key open problem. Diagnostic Query Optimization, a zero-cost query selection rule, adds $+0.120$ to cross-model AUC.