正交化读取是循环记忆中可移除的训练支架
The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory
浏览论文内容
中文总结 AI 辅助
研究发现正交化读取可改善噪声联想记忆,它具有自洽、均匀、可移除特性,能重新调整学习问题。固定预算下解决率衡量突破风险,遵循热/噪声定律。还得出召回基准部分衡量可训练性,系统可检测“涌现”现象的结论。
中文摘要 AI 辅助
最近的一份报告发现,在读取时将mLSTM记忆矩阵正交化(通过五次牛顿-舒尔茨迭代训练)可显著改善噪声联想记忆。该效果可复制,但并非记忆提升。此任务的训练是一个漫长的机会平稳期后紧接着急剧突破,正交化读取通过在平稳期重新调整学习问题起作用。它具有三个特性:自洽性,精确递归最小二乘读取(Mesa层)可重现,而直通减半、增量规则写入、冻结随机键和普通归一化均失败;均匀性,在学习率x硬度网格上,它将突破风险大致提高六倍,且无明显硬度依赖性,拓宽了可行学习率走廊;可移除性,在推理时应用于失败模型无法挽救,按触发突破的时间表退火则使数值上标准的mLSTM保持完全精度。许多已公布的增益根本不需要架构,固定预算下的解决率衡量突破风险,遵循热/噪声定律(学习率弹性+3.0,梯度噪声弹性-1.65),原始词汇量96的结果是大批次噪声条件而非容量条件。直接解码记忆状态表明失败模型约一半的关联以线性可恢复形式存在:平稳期是半写入存储上的读出失败。两个结论超出了干预范围:用于架构选择的召回基准部分衡量可训练性,该系统是“涌现”的完全可检测模型有机体,其中尖锐的行为阈值明显源于对逐渐积累结构的审查指标。
英文摘要
Orthogonalizing the mLSTM memory matrix at read time with five differentiable Newton-Schulz iterations improves noisy associative recall. We replicate this effect and investigate its mechanism. Training on MAD noisy recall exhibits a long chance-level plateau followed by a sharp increase in accuracy. The orthogonalized read improves conditioning during this plateau and can be removed after escape. Ablations support three findings. First, the benefit requires a self-consistent read and gradient: an exact recursive least-squares read (the Mesa layer) yields a similar benefit, while straight-through variants, delta-rule writes, frozen random keys, and Frobenius normalization show no improvement over baseline. Second, across a learning-rate x task-difficulty grid, orthogonalization multiplies escape hazard roughly six-fold, with no detectable dependence on difficulty, and widens the range of learning rates that produce successful runs. Third, adding orthogonalization at inference leaves chance-level failures unresolved, while removing it gradually after escape yields standard mLSTMs at near-perfect accuracy. Schedule changes alone recover much of the reported gain. A batch-size x learning-rate analysis separates the effects of per-step learning rate and gradient noise on escape hazard (elasticities +3.0 and -1.65, respectively), linking the original vocab-96 result to its large-batch training regime. Direct decoding of the memory state recovers roughly half of the associations in behaviorally failed models, indicating a readout-learning limitation despite substantial stored information. These results show that fixed-budget recall benchmarks are sensitive to trainability and provide a tractable setting for investigating abrupt behavioral transitions through measurements of internal representations.
发表机构
- No Way Labs(无路实验室)
机构由 AI 辅助整理,请以论文原文为准。