量子捷径:复相态动力学减少序列模型的优化步数
The Quantum Shortcut: Complex Phase-State Dynamics Reduce the Optimization Steps of Sequence Models
- cAI Technology GmbH(cAI技术有限公司)
- Yam Technology Consulting(雅姆技术咨询公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出将序列模型的基底从实值状态替换为源自量子理论的复值状态,实例化到Mamba和Transformer后,复值模型的优化步数显著少于实值模型,且状态空间模型的优势随训练扩大,注意力模型的优势则衰减至零。
AI中文摘要:
序列模型通常以其骨干网络(即跨位置传递信息的机制,如注意力机制或循环机制)来区分。本文研究的是先于骨干网络、且几乎被所有当前模型共享的一个选择:基底(即表示隐状态的数制,以及从状态到预测的映射形式)。主流基底是具有仿射-softmax读出的实值状态;我们研究一种源自量子理论数学的复值替代方案,其中信息由状态的相位携带,分数为二次玻恩形式。先前工作已证明,这种基底的理想化版本在表示能力上强于任何具有线性读出的实值模型;我们探究它是否也能更快地训练。放宽阻碍部署的两个特性——严格幺正性和玻恩词汇读出——后,我们将其实例化到Mamba状态空间模型和基于注意力的Transformer中。在2.53亿参数(匹配度误差在0.02%以内)、采用同一固定协议在三个字节级语料库上训练的条件下,复值模型达到所有测得的验证损失所需的优化步数约为其实值对应模型的三分之一(状态空间模型)和二分之一(注意力模型)。随后两种骨干网络出现分化:学习率预热结束后,状态空间模型的优势持续扩大,在OpenWebText上的每字符比特数从0.321升至0.354,在FineWeb上从0.368升至0.396,而预热斜坡的人为因素不会导致此现象;注意力模型的优势则在所有语料库上衰减至零,因此是早期训练的效应。
英文摘要:
Sequence models are conventionally distinguished by their backbone, the mechanism that routes information across positions, such as attention or recurrence. This paper varies a choice that is prior to the backbone and shared by nearly all current models: the \emph{substrate}, the number system in which the hidden state is represented together with the form of the map from state to prediction. The prevailing substrate is a real-valued state with an affine--softmax readout; we study a complex-valued alternative drawn from the mathematics of quantum theory, in which information is carried by the phases of the state and scores are quadratic Born forms. Prior work proved an idealized version of this substrate representationally stronger than any real model with a linear readout; we ask whether it also trains faster. Relaxing the two properties that block deployment, exact unitarity and the Born vocabulary readout, we instantiate it in the Mamba state-space model and an attention-based Transformer. At 253M parameters, matched to within $0.02\%$ and trained under one fixed protocol on three byte-level corpora, the complex models reach every measured validation loss in approximately one third (state-space) and one half (attention) of the optimization steps of their real counterparts. The two backbones then diverge. Once the learning-rate warmup ends, the state-space advantage continues to widen, from $0.321$ to $0.354$ bits per character on OpenWebText and from $0.368$ to $0.396$ on FineWeb, which an artifact of the warmup ramp would not do; the attention advantage instead decays toward zero on every corpus, and is therefore an effect of early training.