arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17594cs.LG

当梯度看见秩:训练矩阵记忆中的可证明必要性、因果招募与组合

When the Gradient Sees Rank: Provable Necessity, Causal Recruitment, and Composition in Trained Matrix Memories

  • Pebble ML

机构由 AI 辅助整理,请以论文原文为准。

Samuel Larson

AI总结:

本研究证明梯度训练可学习矩阵记忆所需秩,秩上限决定恢复阈值,有效秩随K增加,并近似理想循环,但大维度下恢复下降。

AI中文摘要:

基于梯度的训练能否学习到在矩阵记忆中存储和组合关联所需的秩?在我们早期的研究中,我们在一个允许秩为1解的任务上使用了矩阵增强推理器,使得该问题悬而未决。我们在$K$个新的键值绑定上训练矩阵记忆,其精确线性恢复需要$\mathrm{rank}(Z) \geq K$。一个固定的线性读出器查询单个矩阵状态,而不访问原始绑定。实验通过余弦相似度大于0.9来衡量恢复,该阈值不同于数学上的相等。学习到的有效秩在测试网格上随$K$增加而增加(在$d = 16$时,Spearman $\rho = 1.0$)。训练时的秩上限在$k = K$附近产生恢复转变:在$d = 8$、$K = 4$时,秩3最多给出0.0004的恢复,而秩4给出0.97。五个种子中有四个在训练算子的21次自应用后仍保持至少0.9996的恢复。在实体子空间上,学习到的算子具有接近$K$的有效秩,并近似理想循环。对于被限制在低于$K$的秩的唯一收敛种子,使用实体子空间算子和理想循环的计算在七次应用内预测测量的余弦值误差在0.008以内。扩展训练解决了若干初始失败,但在编码器宽度固定时,恢复在更大的矩阵维度上仍然下降。

英文摘要:

Can gradient-based training learn the rank needed to store and compose associations in a matrix memory? In our earlier study, we used a matrix-augmented reasoner on a task that admits a rank-1 solution, leaving this question open. We train matrix memories on $K$ fresh key-value bindings whose exact linear recovery requires $\mathrm{rank}(Z) \geq K$. A fixed linear readout queries a single matrix state without access to the original bindings. Experiments measure recovery by cosine similarity greater than 0.9, a threshold distinct from mathematical equality. Learned effective rank increases with $K$ across the tested grid (Spearman $ρ= 1.0$ at $d = 16$). Training-time rank caps produce a recovery transition near $k = K$: at $d = 8$, $K = 4$, rank 3 gives at most 0.0004 recovery and rank 4 gives 0.97. Four of five seeds retain at least 0.9996 recovery through 21-fold self-application of the trained operator. On the entity subspace, the learned operator has effective rank close to $K$ and approximates the ideal cycle. For the single converged seed capped below $K$, a calculation using the entity-subspace operator and ideal cycle predicts the measured cosine within 0.008 through seven applications. Extending training resolves several initial failures, but recovery still declines at larger matrix dimensions with encoder width fixed.

补充信息

↑