arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对数深度循环语言建模

Log-Depth Recurrent Language Modeling

Yiqin Wang, Nuri Cingillioglu, Charles Pert

arXiv 2609.28212首次发表:更新:

发表机构

Imperial College London(帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出将平衡树递归算子扩展至自回归预测,实现对数深度和线性运行时间的语言建模,实验表明其长度外推稳健且性能接近ALiBi Transformer。

AI 中文摘要

尽管Transformer具有固定的计算深度且运行时间随输入token数量呈二次增长,但使用其进行语言建模已变得普遍。另一方面,循环模型提供线性深度但无法并行执行。在本工作中,我们将平衡树递归算子从序列编码扩展到自回归预测,使得所有前缀表示能够以对数深度和线性运行时间计算。我们的实验初步刻画了该模型类别的特性,展示了稳健的长度外推能力以及接近基于ALiBi的Transformer的性能,凸显了其作为语言建模替代架构的潜力。

英文摘要

Language modeling using Transformers has become commonplace despite their fixed computational depth and quadratic runtime with respect to input tokens. Recurrent models on the other hand offer linear depth but no parallel execution. In this work, we extend balanced-tree recursive operators from sequence encoding to autoregressive prediction, enabling all prefix representations to be computed with logarithmic depth and linear runtime. Our experiments provide an initial characterization of this model class, demonstrating robust length extrapolation and performance approaching that of ALiBi-based Transformers, highlighting its potential as an alternative architecture for language modeling.

Comments5 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑