arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一个令牌就够了:从单令牌输出分布中对大语言模型进行指纹识别和验证

One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions

Tomas Bruckner

arXiv 2607.10252首次发表:更新:

发表机构

Faculty of Informatics and Statistics, Prague University of Economics and Business(信息与统计学院,布拉格经济与商业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何从大语言模型单令牌输出分布中进行指纹识别和验证。通过定义行为指纹,利用对简单提示的回答分布,实现模型谱系恢复和验证,还发现生态系统异常,且相关协议等已发布。

AI 中文摘要

大语言模型(LLMs)通过不透明的服务链(API聚合器、经销商和推理提供商)被越来越多地使用,客户无法确认回答的模型是否是宣传的模型,近期审计表明大量商业端点与供应商的参考权重存在偏差。现有识别技术需要长生成文本、令牌级对数概率、对抗性精心设计的提示或模型所有者的合作。我们表明更弱的证据就足够了。我们将LLM的行为指纹定义为其对简单单字提示(如“说出1到100之间的随机数”)的回答的经验分布,以每个查询一个输出令牌的成本在四种语言中收集。通过测量由大型商业聚合器(OpenRouter)提供服务的165个模型,我们发现:(i)这些分布高度非均匀(中位数单元格熵为1.0比特)且特定于模型:同一模型样本的两半比不同模型的样本更接近一个数量级;(ii)指纹之间的 Jensen-Shannon 散度恢复了模型谱系,将模型分配到其记录的家族,留一法准确率为59.5%,而随机率为18.4%;(iii)一种生物特征识别风格的验证协议在完整的40单元格组中实现了7.3%的等错误率,在八个探测单元格中低于11%,每次审计大约一百个单令牌查询。我们还报告了生态系统异常情况,包括一个专有品牌的旗舰端点在分布上与开放权重的Qwen模型无法区分。该协议、提示、原始数据和分析代码已发布以供复制和实际使用。

英文摘要

Large language models (LLMs) are increasingly consumed through opaque serving chains - API aggregators, resellers, and inference providers - in which the client has no technical means to confirm that the model answering is the model advertised, and recent audits show that a substantial fraction of commercial endpoints deviate from the vendor's reference weights. Existing identification techniques require long generated texts, token-level log-probabilities, adversarially crafted prompts, or the model owner's cooperation. We show that far weaker evidence suffices. We define a behavioral fingerprint of an LLM as the empirical distribution of its answers to trivial one-word prompts - "name a random number between 1 and 100" - collected across four languages at a cost of one output token per query. Measuring 165 models served via a large commercial aggregator (OpenRouter), we find that (i) these distributions are highly non-uniform (median cell entropy 1.0 bit) and model-specific: split halves of the same model's samples lie an order of magnitude closer than samples of different models; (ii) Jensen-Shannon divergence between fingerprints recovers model lineage, assigning a model to its documented family with 59.5% leave-one-out accuracy against an 18.4% chance rate; and (iii) a biometric-style verification protocol achieves a 7.3% equal error rate with the full 40-cell battery, and below 11% with eight probe cells - roughly a hundred single-token queries per audit. We further report ecosystem anomalies, including a proprietary-branded flagship endpoint distributionally indistinguishable from an open-weight Qwen model. The protocol, prompts, raw data, and analysis code are released for reproduction and operational use.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑