发表机构
Pebble ML(Pebble ML)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现Matrix-CODI模型在ProsQA任务中秩k投影消融曲线平坦,梯度对Z的秩不敏感,排除读出方式影响后仍存在该现象,普通GPT-2 SFT也有类似表现,揭示秩盲与位置无关性的混淆问题。
AI 中文摘要
连续思维链模型将推理压缩为隐式token,矩阵变体通过d×d矩阵瓶颈路由每个隐式token,引入秩作为隐式矩阵Z的单样本结构可观测值。若矩阵隐式通过叠加携带并行推理路径,秩应跟踪这些路径,将Z截断为低秩会损害需多组件的任务准确率。在Matrix-CODI模型的四种训练 regime(三种针对ProsQA,一种针对低于学习阈值的GSM8K-Aug)中,秩k投影消融曲线的波动在0.6个百分点以内。三次随机种子重复实验的准确率为81.0±2.0个百分点,而Z的最终有效秩在{4,12,13}之间,损失未奖励任何特定秩。为测试秩盲是否仅源于“展平后投影”读出方式,我们训练了四种读出方式:双线性重参数化、对Z非线性的双线性加GELU读出、通过MLL馈送奇异值的SVD增强读出、对ZZ^T的二次读出。所有四种秩k曲线均保持平坦(斯皮尔曼p值分别为0.63、0.14、0.82、0.46),且对Z非线性的读出仍保持平坦曲线。对Z的线性探测在目标预测上表现弱于原始预训练隐状态(AUC为0.673 vs 0.846)。对普通GPT-2 SFT(无矩阵瓶颈、无Z,三次随机种子,n=500)的阴性对照在相同干预范式下重现了平坦的秩k曲线,合并均值范围为0.20个百分点,随机h敏感度下限达到相同准确率:秩k消融本身将秩盲与位置无关性混淆了。
英文摘要
Continuous chain-of-thought models compress reasoning into latent tokens. Matrix-valued variants, which route each latent token through a d x d matrix bottleneck, introduce rank as a single-sample structural observable on the latent matrix Z. If matrix latents carry parallel reasoning paths via superposition, rank should track them, and truncating Z to low rank should hurt accuracy on tasks whose solutions plausibly require multiple components. Across four training regimes of a matrix-CODI model (three on ProsQA, one on GSM8K-Aug below the learning threshold), the rank-k projection ablation curve is flat to within 0.6 percentage points. A three-seed replication yields 81.0 +/- 2.0 percentage points accuracy while the final effective rank of Z spans {4, 12, 13}; the loss does not reward any particular rank. To test whether rank-blindness arises from the flatten-then-project readout alone, we trained four readouts: a bilinear reparametrization, a bilinear-plus-GELU readout nonlinear in Z, an SVD-augmented readout feeding singular values through an MLP, and a quadratic readout in Z Z^T. All four rank-k curves remain flat (Spearman p-values 0.63, 0.14, 0.82, 0.46). The flat curves persist for readouts nonlinear in Z. A linear probe on Z underperforms a raw pretrained hidden state at target prediction (AUC 0.673 vs. 0.846). A negative control on vanilla GPT-2 SFT (no matrix bottleneck, no Z, three seeds, n=500) reproduces a flat rank-k curve under the same intervention paradigm with pooled-mean range 0.20pp, and a random-h sensitivity floor lands at the same accuracy: the rank-k ablation alone conflates rank-blindness with position-irrelevance.
CommentsAccepted at the ICML 2026 Mechanistic Interpretability Workshop, https://openreview.net/forum?id=Spof4PusVI. 9 pages. Corrects a data-entry error in the workshop version: the seed-1337 accuracy in the three-seed replication was reported as 80.47% (a control run); the archived value is 78.91%, so the three-seed mean is 81.0 +/- 2.0pp (was 81.5 +/- 1.2pp). All other results are unchanged