arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有自适应退出状态选择的循环状态空间语言模型

CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models

Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa

arXiv 2607.10110首次发表:更新:

发表机构

Rikkyo University; The University of Tokyo(立教大学; 东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究循环状态空间语言模型在推理任务及预训练中的表现,采用循环曼巴等架构,在推理任务中优于非循环基线,预训练时在下游基准测试有竞争力,适配退出门后在中间深度提升下游性能。

AI 中文摘要

近期关于循环语言模型的研究表明,许多推理问题受益于更大的计算深度而非更多独立参数。现有研究几乎只关注Transformer架构,未探讨该原理是否适用于状态空间语言模型。本文研究了循环曼巴和循环混合曼巴 - Transformer架构,在两个推理任务上,循环曼巴优于参数匹配的非循环基线。在等参数和等FLOPs协议下预训练语言模型时,循环模型在下游基准测试中具有竞争力。最后,为循环曼巴适配Ouro的两阶段退出门,自适应退出状态选择在中间深度提高了下游性能。

英文摘要

Recent work on looped language models suggests that many reasoning problems benefit from greater computational depth rather than from additional independent parameters. Existing studies, however, focus almost exclusively on Transformer backbones, leaving open whether this principle also applies to state-space language models. We investigate Looped Mamba and Looped Hybrid Mamba-Transformer architectures, which repeatedly apply a shared Mamba (or hybrid) block to introduce explicit finite-depth recurrent computation. On two controlled reasoning tasks-Mano (modular-arithmetic manipulation) and p-hop induction-Looped Mamba consistently outperforms parameter-matched non-looped baselines and, in several settings, matches or exceeds non-looped models of equal effective depth. We then extend the study to language model pre-training under matched iso-parameter and iso-FLOPs protocols, which jointly disentangle the effects of parameter sharing and effective depth: looped models remain competitive on downstream benchmarks with substantially fewer distinct parameters, although deeper non-looped models retain an advantage in validation perplexity under strict iso-FLOPs comparisons. Finally, we adapt Ouro's two-stage exit gate to Looped Mamba for threshold-controlled selection among recurrent-step outputs. Executing such exits on a state-space backbone, however, leaves the recurrent state without its deeper updates, and validation perplexity then degrades severely. We therefore introduce a cache-hole adaptation that aligns continued training with skipped-state inference. At the scales studied, the adapted model keeps perplexity close to full computation and matches or exceeds full-compute exit-state selection on downstream benchmarks while executing roughly half of the recurrent steps, which translates into measured inference speedups once the prefill is compute-bound.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑