AI 中文总结
该研究在Qwen2.5-0.5B-Instruct中融入循环深度,在两个参数预算下实现安装,其模型在ARC测试集等任务上性能优于同等规模基准,可外推至1.5倍监督深度,推理速度更快。
AI 中文摘要
一个稠密的预训练语言模型可被改造融入循环深度,学习一种仅基于结果的退火后仍能保留的迭代隐式转换机制。Qwen2.5-0.5B-Instruct被拆分为Prelude(前序模块)、权重绑定的循环块和Coda(后序模块),其中包含一条保持恒等性的单循环路径以及后续循环的重入桥。在第1个循环时,该改造模型在预注册的ARC测试集上的性能不劣于其基础模型。研究得出三项发现:其一,该机制是可复用的过程而非终端答案查找,且可在两个参数预算下安装:冻结基础权重时需600万训练参数,完整块则需1.8亿参数。在中间步骤监督下,模型每个循环计算一个任务步骤,仅对最终答案评分时仍能保持性能;适配器整体性能与完整块相当(83.8%对比84.0%),可支持至第11个循环,超出后性能略有落后。在受控的文本表述上,文本微调达到79%-86%(零样本迁移效果极弱),从已安装机制启动的适配器文本训练比匹配的全新训练高出18.6个百分点,在保留测试集上也表现更佳。其二,该操作可外推至约1.5倍的监督深度,在第18个循环时仍保持70%的准确率。其三,相同规模的临时训练模型在其学习范围内与循环模型性能相当,但超出后性能崩溃;循环模型整体胜出,准确率为84%对比72%,在第10个循环后保留53%对比2.5%,推理速度快7.6倍。因此,迭代Transformer在隐空间中可执行比同等或更大规模的微调模型更快的深度推理,这是系统级对比的结果。第二项任务(反向运行规则)暴露了局限性:反向规则可单独学习,但无法在保留已安装机制和通用能力的同时实现延续,存在灾难性干扰边界;学习深度选择仍是未解决的问题。
英文摘要
A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after outcome-only annealing. Qwen2.5-0.5B-Instruct is split into a Prelude, a weight-tied Recurrent Block, and a Coda, with an identity-preserving one-loop path and a re-entry bridge on later loops. At loop 1 the retrofit remains non-inferior to its base on a preregistered ARC battery. Three findings. First, the mechanism is a reusable procedure rather than terminal-answer lookup, and installs at two budgets: 6M trained parameters over frozen base weights and 180M full-block. With intermediate-step supervision, the model computes one task step per loop and persists when only final answers are graded. The adapter matched the full block overall (83.8% versus 84.0%), led through depth 11, and trailed beyond. Verbal fine-tuning reached 79-86% on controlled verbal renderings (zero-shot transfer was minimal), and adapter verbal training begun from the installed mechanism outpaced matched fresh training by 18.6 points, including on a held-out test set. Second, the operation extrapolates to roughly 1.5 times its supervised depth, holding 70% accuracy through depth 18. Third, a same-size scratchpad-trained model matched the recurrent model within its learned horizon but collapsed beyond it. The recurrent model won overall, 84% versus 72%, retained 53% versus 2.5% beyond depth 10, and answered 7.6 times faster. An iterative transformer can therefore perform deeper reasoning in latent space faster than comparable or larger models fine-tuned on the same task, in a system-level comparison. A second task, running the rule in reverse, exposed the limits: the inverse was learnable in isolation, but no continuation acquired it while preserving the installed mechanism and general capability, a catastrophic-interference boundary. Learned depth selection remains open.