arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型迭代自馈探测的测量对象,以及一种区分构造与模型的测试方法

What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model

Nicolás Vera Zúñiga

arXiv 2608.10986首次发表:更新:

AI 中文总结

本研究探究语言模型迭代自馈探测的测量对象,提出区分探测构造与模型的测试方法,揭示两类测量量的差异,复现验证工具并纠正错误判断。

AI 中文摘要

一类日益增多的方法通过向语言模型馈送自身输出来探测该模型,包括自一致性、迭代精化、智能体循环等。我们探究此类探测的测量对象,为此选择一个构造来使该问题变得尖锐:一个由 token 单元构成的环,通过模型自身的窗口条件概率 p_r(x_i | x_{i±r}) 对其进行原位重采样。其底层机制是 token 序列上的 Glauber 动力学,并非新内容;我们所改变的是耦合方式。在共同随机数下推进仅相差一个 token 的两个环,会使未受损副本的差异恰好为零,因此损伤传播在最大耦合给出混合时间的地方变得可测量。答案是,它在外观相似的读数中同时测量两种不同的事物。部分量由构造固定:损伤光锥是运动学的,且 token 空间李雅普诺夫指数 λ_ca(r) 的半径缩放在 19 个模型和跨越 70 倍的两个尺度阶梯上具有模型不变性。其他量则真实地跟踪模型:λ_ca 在训练中一个可复现的点处穿过零点,且吸引子份额无论晶格如何构建都能一致地对模型进行排名。若不加以区分,第一种量极易被误认为是第二种——我们自己就犯了这样的错误,持续了四个月,还报告了一个我们测量到三位小数的相变,该相变属于探测而非任何语言模型。我们给出了区分它们的测试方法:保持构造固定并改变模型,或保持模型固定并改变构造,观察哪些读数发生变化。我们首先通过复现验证了该工具,与独立预测完全一致地恢复了 Domany-Kinzel 损伤场,还报告了该方法发现的估计器故障——四个被撤回的判断,每个都涉及一个看似测量值的量。该方法论作为一个包提供。

英文摘要

A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. We ask what such a probe measures, in a construction chosen to make the question sharp: a ring of token cells resampled in place by the model's own windowed conditional p_r(x_i | x_{i+-r}). The substrate is Glauber dynamics on token sequences and is not new; what we change is the coupling. Advancing two rings that differ in one token under common random numbers makes undamaged copies diverge by exactly zero, so damage spreading becomes measurable where a maximal coupling gives mixing times instead. The answer is that it measures two different things at once, in readings that look alike. Some quantities are fixed by the construction: the damage light cone is kinematic, and the radius scaling of the token-space Lyapunov exponent lambda_ca(r) is model-invariant across 19 models and two scale ladders spanning 70x. Others genuinely track the model: lambda_ca crosses zero at a reproducible point in training, and the attractor share ranks models consistently however the lattice is built. Left undistinguished, the first kind is readily mistaken for the second -- we did so ourselves for four months, and report a phase transition we measured to three decimal places that belongs to the probe rather than to any language model. We give the test that separates them: hold the construction fixed and vary the model, or hold the model fixed and vary the construction, and see which readings move. We validate the instrument by reproduction first, recovering a Domany-Kinzel damage field bit-exactly against an independent prediction, and we report the estimator failures that this discipline caught -- four retracted verdicts, each on a quantity that looked like a measurement. The methodology ships as a package.

Comments16 pages, 4 figures. Code, per-run results, and the findings ledger: https://github.com/nicoveraz/token-lattice-ca (archived: https://doi.org/10.5281/zenodo.21880472)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑