发表机构
Singapore Institute of Technology; IIIT Delhi; Nanyang Technological University; NVIDIA AI Technology Centre(新加坡理工学院; 德里印度信息技术学院; 南洋理工大学; 英伟达人工智能技术中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出分层局部注意力的TextNCA模型作为分析探针,发现其性能受从窄到宽的分阶段调度主导,迭代和GRU门控等仅带来小增益。
AI 中文摘要
严格局部、迭代、权重共享的计算基元能否支撑语言建模,且这三个属性中哪一个实际驱动模型行为?我们定义了TextNCA,即神经细胞自动机基元的一维因果窗口注意力实现,并研究了其分层变体,该变体在WikiText-103数据集上采用三个级联阶段,窗口大小w∈{8,32,128},每个阶段有Tₛ个权重共享迭代,模型规模约30M参数,训练步数为60k。该规模下模型未达到参数匹配的Transformer性能(Hier-TextNCA困惑度(PPL)为60.3,Transformer-6L为52.8,Transformer-12L为44.7),因此我们将其视为分析探针而非替代方案。观察到的行为主要由从窄到宽的分阶段调度解释:采用相同调度的非迭代滑动窗口Transformer与迭代模型的PPL仅相差+4.1,而反转、扁平化或破坏调度的单调顺序会使PPL增加16.7至70.8。迭代在调度基础上带来较小的有界收益,在Tₛ=4时达到明显最优,超过该值后性能呈U型退化。GRU门控和学习到的每步嵌入是该收益出现的必要条件,使用随机Tₛ训练可在推理时获得迭代次数调节旋钮,但代价是绝对PPL大幅升高。我们将本研究定位为对NCA式计算各部分在语言建模中权重的可控分析。
英文摘要
Can a strictly local, iterated, weight-shared computation primitive support language modelling, and which of those three properties actually drives the model's behaviour? We define \textsc{TextNCA}, a 1D causal windowed-attention realisation of the Neural Cellular Automaton primitive, and study a hierarchical variant that cascades three stages with windows $w \in \{8, 32, 128\}$ and $T_s$ shared-weight iterations per stage, all on WikiText-103 at roughly 30M parameters and 60k training steps. The model does not match a parameter-matched Transformer at this scale (Hier-TextNCA $60.3$ vs.\ Transformer-6L $52.8$ and Transformer-12L $44.7$ PPL), so we treat it as an analytical probe rather than a proposed alternative. The behaviour we observe is largely explained by the staged narrow-to-wide schedule: a non-iterating sliding-window Transformer that reuses the same schedule comes within $+4.1$ PPL of the iterated model, while reversing, flattening, or breaking the monotonic ordering of the schedule costs between $+16.7$ and $+70.8$ PPL. Iteration adds a smaller bounded benefit on top of the schedule, with a clear optimum at $T_s{=}4$ and a U-shaped degradation beyond it. The GRU gate and learned per-step embeddings are required for that benefit to appear, and training with random $T_s$ yields an inference-time iteration-count knob at the cost of substantially higher absolute PPL. We position the work as a controlled reading of which parts of NCA-style computation carry the weight in language modelling.