行为是推理发展的不完整衡量指标:循环深度推理器的跨表面预到达可访问性与发展推理的局限性
Behaviour Is an Incomplete Measure of Reasoning Development: Cross-surface pre-arrival accessibility and the limits of developmental inference in a recurrent-depth reasoner
- Prime Calibre
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究以30M参数的循环深度关系推理器为对象,发现行为能力、内部可访问性与训练发展是不同指标,行为无法完整衡量推理发展,需因果干预进一步研究。
AI中文摘要:
能力发展通常通过行为阈值、最终检查点或解码器能从隐藏状态中读取的内容来推断,但这些量未必对应同一事件。我们在一个由预言机定义的封闭世界中研究了一个30M参数的循环深度关系推理器,全程将训练时间轴与推理时间轴分开,采用密集行为轨迹、两个训练表面、预先注册的预到达隐藏状态探针、前瞻性可评估性检查以及明确的未训练和阴性对照。首先看行为:在一个冻结的获取标准下,三跳能力在符号表面上花费了70个逻辑周期,在语言表面上则花费了13055个周期,两者相差186.5倍,之后语言表面上的四跳能力在8个逻辑周期内完成。在13055个周期的训练过程中,四跳的保留行为从未超过3/40,最终为0/40。接着看内部测量:在语言表面上,线性探针在行为到达前恢复了未来答案的身份,准确率为0.056159,而均匀随机概率为0.025,未训练对照为0.024758,总体频率基线为0.048309(p=0.012987;涉及16/40个答案类别)。类似的预到达可访问性在表面变化后仍然存在,在上游结构位置达到0.1020,零步对照为0.0460(p=0.000999),在读出比较器处为0.0618(p=0.004),涉及21/40个类别。最后,尝试在训练过程中追踪这种可访问性的做法无法清晰评估:探针资格由行为到达定义,因此测量的总体随被测量对象而变化。行为能力、内部可访问性和训练时间发展是不同的可观测指标,行为和解码器可访问性均无法识别训练获得的计算;因果干预是必要的下一步。
英文摘要:
Capability development is routinely inferred from behavioural thresholds, from final checkpoints, or from what a decoder can read out of a hidden state. These quantities need not identify the same event. We study a 30M-parameter recurrent-depth relational reasoner in a closed, oracle-defined world, using dense behavioural trajectories, two training surfaces, preregistered pre-arrival hidden-state probes, prospectively checked evaluability, and explicit untrained and negative controls, holding the training-time and inference-time axes separate throughout. Behaviour first: under one frozen acquisition criterion, three-hop competence cost 70 logical epochs on the symbolic surface and 13,055 on the verbal surface, a 186.5-fold contrast, after which verbal four-hop competence cleared in 8 logical epochs. Across the 13,055-epoch grind, four-hop held-out behaviour never exceeded 3/40 and ended at 0/40. Internal measurement next: on the verbal surface a linear probe recovered future-answer identity before behavioural arrival at 0.056159 against uniform chance 0.025, an untrained control of 0.024758 and a population frequency baseline of 0.048309 (p = 0.012987; 16/40 answer classes contributing). Analogous pre-arrival accessibility survived the surface change, reaching 0.1020 against a zero-step control of 0.0460 (p = 0.000999) at the upstream structural position and 0.0618 at the readout comparator (p = 0.004), with 21/40 classes contributing. Finally, the natural attempt to track that accessibility across training was not cleanly evaluable: probe eligibility is defined by behavioural arrival, so the measured population changes with the measurand. Behavioural competence, internal accessibility, and training-time development are distinct observables, and neither behaviour nor decoder accessibility identifies the computation training acquired; causal intervention is the necessary next step.