发表机构
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过24对端点的冻结阈值保留实验,验证了Token计数一致性可作为共享分词栈的指纹,但否定其作为模型家族谱系独立必要测试的用途。
AI 中文摘要
当大语言模型(LLM)通过中继和经销商API提供服务时,黑盒模型归因变得愈发重要。一个颇具吸引力的低成本信号是OpenAI兼容端点返回的提示Token计数:两个共享分词器和聊天模板的模型,其计数序列可能在固定偏移范围内一致。然而,该信号在更广泛的模型家族归因中的有效性,尚未得到直接的保留测试验证。我们针对24对带标签的端点开展冻结阈值研究,这些端点对半划分为开发集和未触及的保留集,包含3次时间重复,每对对应30篇受控文本。我们引入了有效性门控结果契约,以区分观测到的不相似性与因缺失使用数据、速率限制或端点策略导致的无信息测量。所得的平移不变精确匹配得分完美区分了12对开发集,得出冻结阈值为0.725。但在保留集上,按预先指定的三次重复规则,仅12对中的6对符合条件。在符合条件的对中,平衡准确率为0.75,灵敏度为0.50(95% Wilson区间0.15--0.85),特异度为1.00(0.342--1.00)。两对同家族模型——Qwen 3.8和DeepSeek V4变体——低于冻结阈值。在4320次正式API调用中,所有日志均可重放,而保留集包含189个非200响应和157个无提示Token使用的成功响应。因此,本研究验证了Token计数一致性可作为共享分词栈的指纹,但否定了其作为模型家族谱系独立必要测试的用途。
英文摘要
Black-box model attribution is increasingly relevant when large language models (LLMs) are served through relay and reseller APIs. A tempting low-cost signal is the prompt-token count returned by an OpenAI-compatible endpoint: two models that share a tokenizer and chat template may produce the same count sequence up to a fixed offset. Yet the validity of this signal for broader \emph{model-family} attribution has received little direct holdout testing. We conduct a frozen-threshold study over 24 labeled endpoint pairs, split evenly into a development set and an untouched holdout set, with three temporal repeats and 30 controlled texts per pair. We introduce a validity-gated result contract that distinguishes an observed dissimilarity from an uninformative measurement caused by missing usage data, rate limits, or endpoint policy. The resulting shift-invariant exact-match score perfectly separates the 12 development pairs, yielding a frozen threshold of 0.725. On holdout, however, only 6 of 12 pairs are eligible under the pre-specified three-repeat rule. Among eligible pairs, balanced accuracy is 0.75, sensitivity is 0.50 (95\% Wilson interval 0.15--0.85), and specificity is 1.00 (0.342--1.00). Two same-family pairs---Qwen 3.8 and DeepSeek V4 variants---fall below the frozen threshold. Across 4,320 formal API calls, every log is replayable, while holdout contains 189 non-200 responses and 157 successful responses without prompt-token usage. The study therefore validates token-count consistency as a fingerprint of a shared \emph{tokenization stack}, but rejects its use as a standalone necessary test for model-family lineage.
Comments10 pages, 2 figures