ZetaGPT:无位置编码的状态空间注意力语言模型的参考实现
ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models
浏览论文内容
中文总结 AI 辅助
ZetaGPT是首款无显式位置编码的开源小型语言模型,它将因果状态空间方程与自注意力结合,提供全开源训练流水线,为无位置编码语言模型的研究提供紧凑可复现的参考实现。
中文摘要 AI 辅助
基于Transformer的语言模型依赖自注意力机制,其计算具有排列等变性,因此缺乏表示令牌顺序的内在机制。现有架构通过显式引入位置信息(如学习到的位置嵌入或手工设计的位置编码,例如旋转位置编码RoPE)来解决这一局限,将位置信息视为架构获取的能力而非模型的固有属性。受追求无位置编码架构的驱动,本研究探索一种语言模型架构,该架构在注意力计算前集成因果状态空间方程以隐式编码位置信息。具体而言,每个模型块在自注意力前应用因果状态空间方程,使循环状态动力学将序列信息编码到令牌表示中。因此,后续注意力层可基于感知位置的表示运行,无需显式位置编码,同时保留自注意力的表达建模能力。我们推出ZetaGPT,一款面向研究、快速原型开发、算法验证及教育应用的紧凑混合语言模型。除提出的架构外,ZetaGPT还提供完全开源的端到端训练流水线,涵盖数据集构建、分词器训练、预训练、监督微调、基于人类反馈的强化学习(RLHF)以及通过纯强化学习实现的思维链(CoT)推理。据我们所知,ZetaGPT是首款无显式位置编码的开源小型语言模型,为无位置编码语言模型的开发与实证研究建立了紧凑、可复现的参考实现。
英文摘要
Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mechanism for representing token order. Existing architectures address this limitation by explicitly incorporating positional information through learned positional embeddings or hand-crafted positional encodings, such as rotary positional encoding (RoPE), treating positional information as an architecturally acquired capability rather than an inherent property of the model. Motivated by the pursuit of positional-encoding-free architectures, this work explores a language model architecture that integrates causal state-space equations to implicitly encode positional information before attention computation. Specifically, each model block applies a causal state-space equation before self-attention, allowing recurrent state dynamics to encode sequential information into token representations. Consequently, subsequent attention layers operate on position-aware representations without requiring explicit positional encodings while retaining the expressive modeling capacity of self-attention. We present \textsc{ZetaGPT}, a compact hybrid language model designed for research, rapid prototyping, algorithm verification, and educational applications. In addition to the proposed architecture, \textsc{ZetaGPT} provides a fully open-source, end-to-end training pipeline encompassing dataset construction, tokenizer training, pretraining, supervised fine-tuning, reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) reasoning via pure reinforcement learning. To the best of our knowledge, \textsc{ZetaGPT} is the first open-source small language model without explicit positional encoding and establishes a compact, reproducible reference implementation for the development and empirical study of positional-encoding-free language models.
发表机构
- University of Galway(戈尔韦大学)
机构由 AI 辅助整理,请以论文原文为准。