arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31181cs.CL

模型将其自身重复标记发送至何处

Where a Model Sends Its Own Repeated Token

Nicolás Vera Zúñiga

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过分析模型对自身重复标记的预测映射,提出一种黑盒模型识别方法,发现非固定点映射能有效区分模型家族,且优于分词器特征,并验证了其鲁棒性。

中文摘要 AI 辅助

黑盒模型识别通过评分模型对自然语言提示的响应来工作。一条研究路线向模型提供退化输入——其自身的重复标记——以寻找失败模式而非身份。我们采用该输入,并询问当模型不这样做时它去向何处。对于每个标记 t,在一次前向传递中读取 argmax p(. | t, t);结果是对整个词汇表的映射,分为两部分。第一部分——哪些标记是固定点——部分被预期到,我们将其报告为失败的估计量:其上的自然距离为 83% 的基数,在 3471 个样本中以零精度下限将语料库操纵分离两个比特,并以 0.5833 的属性归因于家族。第二部分,即映射将非固定点标记发送至何处,未被记录;唯一持有这些标记的论文将其记录为零。基于源标记的配对通过构造消除了基数混淆(r 从 0.9128 降至 -0.0932),并以 0.8333 的属性归因于家族——在十九个模型的池中评分十二个模型——机会为 0.1389,跨越七个分词器组和多个语料库。两个零假设清除它:频率匹配的目的地一致率为 0.1429,独立边际为 0.0798。家族预测一致性优于分词器(0.2031 对 0.1205),且循环架构以平衡准确率 1.0 聚类,对比多数类率 0.7895,或在排除每个模型的主导目的地后为 0.90——这是我们坚持的数字。我们测量鲁棒性包络:8 位权重舍入对映射的影响小于对训练语料库去重的影响(在一个支持集上为 0.9004 对 0.6353),4 位则破坏它(0.0098;在部署粒度下为 0.1812,因此不是粗粒度伪影),且精度下限因模型而异,从 0.201 到 0.9778。所有估计量和终止条件在数据之前注册,失败的那个与存活的那个一样被完整报告。

英文摘要

Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We take that input and ask where the model goes when it does not. For each token t, read argmax p(. | t, t) in one forward pass; the result is a map on the whole vocabulary, with two halves. The first -- which tokens are fixed points -- is partially anticipated, and we report it as a failed estimand: the natural distance on it is 83% cardinality, separates a corpus manipulation by two bits in 3471 against a precision floor of zero, and attributes families at 0.5833. The second half, where the map sends tokens that are not fixed points, is unrecorded; the one paper holding those tokens logged them as a zero. Pairing on the source token removes the cardinality confound by construction (r from 0.9128 to -0.0932) and attributes families at 0.8333 -- twelve models scored against a pool of nineteen -- with chance 0.1389, across seven tokenizer groups and several corpora. Two nulls clear it: frequency-matched destinations agree at 0.1429, independent marginals at 0.0798. Family predicts agreement better than tokenizer (0.2031 against 0.1205), and recurrent architectures cluster at balanced accuracy 1.0 against a 0.7895 majority rate, or 0.90 once each model's dominant destination is excluded -- the figure we stand behind. We measure the robustness envelope: 8-bit weight rounding moves the map less than deduplicating the training corpus does (0.9004 against 0.6353, on one support), 4-bit destroys it (0.0098; 0.1812 at deployment granularity, so not a coarseness artefact), and the precision floor varies by model from 0.201 to 0.9778. All estimands and kill conditions were registered before the data, and the failed one is reported as fully as the surviving one.

发表机构

  • Independent Researcher(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑