AdaptPrint:黑盒大语言模型服务的响应自适应指纹识别
AdaptPrint: Response-Adaptive Fingerprinting of Black-Box LLM Services
浏览论文内容
中文总结 AI 辅助
针对黑盒LLM服务的不透明性问题,提出响应自适应指纹识别方法AdaptPrint,整合三种探测策略实现LLM身份识别,在27个候选模型上优于现有方法,准确率达80.6%等且鲁棒性强。
中文摘要 AI 辅助
黑盒大语言模型(LLM)服务已成为实用的部署范式,但其不透明性阻碍了安全风险的系统评估,也增加了模型所有者的版权审计难度。黑盒LLM指纹识别通过查询-响应交互识别底层LLM身份,为弥合这一差距提供了有前景的途径。现有方法通常使用固定查询集从目标LLM服务收集响应,在面对现实复杂配置(如系统提示和采样设置)时表现不佳。为克服这些局限,我们提出AdaptPrint,一种用于揭示黑盒LLM服务中隐藏LLM身份的响应自适应指纹识别方法。AdaptPrint整合了三种渐进式响应一致性探测策略:直接探测、续接探测和跟进探测。AdaptPrint通过在候选LLM间进行相似性匹配来确定最终LLM身份。实验结果显示,在27个候选模型中,AdaptPrint显著优于最先进的方法,达到Top-1、Top-3和Top-5准确率分别为80.6%、90.3%和92.1%,且在不同防御策略和解码参数下表现出强鲁棒性。
英文摘要
Black-box LLM services have emerged as a practical deployment paradigm. Nevertheless, their opacity also hinders the systematic assessment of security risks and complicates copyright auditing for model owners. Black-box LLM fingerprinting, which identifies the underlying LLM identity through query-response interactions, offers a promising way to bridge this gap. Existing approaches typically collect responses from target LLM services using a fixed set of queries and perform poorly in the presence of realistic and complex configurations (e.g., system prompt and sampling settings). To overcome these limitations, we propose AdaptPrint, a response-adaptive fingerprinting method for revealing hidden LLM identities in black-box LLM services. AdaptPrint integrates three progressive response consistency probing strategies: Direct Probing, Continuation Probing, and Follow-up Probing. AdaptPrint determines the final LLM identity by performing similarity matching among candidate LLMs. Experimental results show that AdaptPrint significantly outperforms state-of-the-art methods among 27 candidate models, achieving Top-1, Top-3, and Top-5 accuracies of 80.6%, 90.3%, and 92.1%. AdaptPrint also demonstrates strong robustness across different defense strategies and decoding parameters.