arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TokenPrint:用于语言模型溯源的校准标记空间指纹

TokenPrint: A Calibrated Token-Space Fingerprint for Language-Model Provenance

Yuqi Wu, Shengming Zhao, Jie Chen

arXiv 2608.08139首次发表:更新:

发表机构

College of Biomedical Engineering, Fudan University(复旦大学生物医学工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出无需训练的语言模型溯源指纹TokenPrint,基于250个固定知识探针的隐藏状态投影,可有效检索谱系、区分模型相关性,且对量化稳定,在32个开放权重模型上验证了其性能。

AI 中文摘要

确定语言模型的溯源信息,包括其基础检查点和训练分布中可能存在的重叠,是仅靠元数据无法解决的治理挑战。我们引入一种无需训练的指纹,该指纹基于250个固定知识探针所引出的后期隐藏状态的前k个词汇投影,通过解码后的标记字符串的Jaccard重叠进行比较。我们在来自9个家族的32个开放权重模型(规模从0.6B到32B)上对该方法进行评估,这些模型具有已记录的关系:(1)一个“相似度阶梯”大致遵循模型的相关性:在相同数据上独立训练的模型的原始得分为0.48(词汇校正后为0.35),其次是共享基础的微调模型(0.39/0.33)、同一开发者的相关模型(0.38/0.28),以及无已记录关系的模型(0.22/0.17)。这种相同数据的信号在三个机构、两个分词器家族和两个架构类别中均存在,且在训练的前1%即可出现,早于可测量的任务能力,表明共享训练数据的贡献超出了能力收敛的范畴。(2)作为一种最近邻“谱系检索”方法,该指纹在所有5次R1蒸馏中均将确切记录的基础模型排在前两位候选中(平均排名1.8,MRR为0.60),其中包括无法通过粗略元数据识别的数学专用基础模型。(3)“深度消融”实验表明,谱系组的区分度向输出分布增强,AUC从四分之一深度的0.72上升到输出层的0.90;仅使用前5个输出标记仍能保持0.87的AUC。(4)该指纹在量化下保持稳定,int8量化下的Jaccard相似度为0.92,int4量化下为0.82至0.85,而校准池中的最大跨模型相似度为0.81。我们发布了这些探针、代码和指纹。

英文摘要

Establishing the provenance of a language model---including its base checkpoint and possible overlap in training distributions---is a governance challenge that metadata alone cannot resolve. We introduce a training-free fingerprint based on the top-$k$ vocabulary projections of late hidden states elicited by 250 fixed knowledge probes, compared using Jaccard overlap over decoded token strings. We evaluate the method on 32 open-weight models from nine families (0.6B--32B) with documented relationships. (1)~A \emph{similarity ladder} broadly follows model relatedness: independently trained models on identical data score 0.48 raw (0.35 vocabulary-corrected), followed by shared-base fine-tunes (0.39/0.33), same-developer relatives (0.38/0.28), and models with no documented relationship (0.22/0.17). This identical-data signal persists across three organizations, two tokenizer families, and two architecture classes, and emerges within the first 1\% of training before measurable task competence, suggesting a contribution from shared training data beyond capability convergence. (2)~As a nearest-neighbor \emph{lineage-retrieval} method, the fingerprint ranks the exact documented base among the top two candidates for all five R1 distillations (mean rank 1.8, MRR 0.60), including a math-specialized base not identifiable from coarse metadata. (3)~A \emph{depth ablation} shows that lineage group discrimination strengthens toward the output distribution, with AUC increasing from 0.72 at quarter depth to 0.90 at the output; using only the top 5 output tokens retains AUC 0.87. (4)~The fingerprint remains stable under quantization, with Jaccard similarity of 0.92 under int8 and 0.82--0.85 under int4, compared with a maximum cross-model similarity of 0.81 in the calibration pool. We release the probes, code, and fingerprints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑