arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TextNCA:通过分层局部注意力实现语言建模的神经细胞自动机

TextNCA: Neural Cellular Automata for Language Modeling via Hierarchical Local Attention

Avni Mittal, Avinash Anand, Ashutosh Kumar, Dikshant Kukreja, Kritarth Prasad, Sushane Dulloo, Erik Cambria, Timothy Liu, Zhengkui Wang, Rajiv Ratn Shah

arXiv 2608.02050首次发表:更新:

发表机构

Singapore Institute of Technology; IIIT Delhi; Nanyang Technological University; NVIDIA AI Technology Centre(新加坡理工学院; 德里印度信息技术学院; 南洋理工大学; 英伟达人工智能技术中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出分层局部注意力的TextNCA模型作为分析探针,发现其性能受从窄到宽的分阶段调度主导,迭代和GRU门控等仅带来小增益。

AI 中文摘要

严格局部、迭代、权重共享的计算基元能否支撑语言建模,且这三个属性中哪一个实际驱动模型行为?我们定义了TextNCA,即神经细胞自动机基元的一维因果窗口注意力实现,并研究了其分层变体,该变体在WikiText-103数据集上采用三个级联阶段,窗口大小w∈{8,32,128},每个阶段有Tₛ个权重共享迭代,模型规模约30M参数,训练步数为60k。该规模下模型未达到参数匹配的Transformer性能(Hier-TextNCA困惑度(PPL)为60.3,Transformer-6L为52.8,Transformer-12L为44.7),因此我们将其视为分析探针而非替代方案。观察到的行为主要由从窄到宽的分阶段调度解释:采用相同调度的非迭代滑动窗口Transformer与迭代模型的PPL仅相差+4.1,而反转、扁平化或破坏调度的单调顺序会使PPL增加16.7至70.8。迭代在调度基础上带来较小的有界收益,在Tₛ=4时达到明显最优,超过该值后性能呈U型退化。GRU门控和学习到的每步嵌入是该收益出现的必要条件,使用随机Tₛ训练可在推理时获得迭代次数调节旋钮,但代价是绝对PPL大幅升高。我们将本研究定位为对NCA式计算各部分在语言建模中权重的可控分析。

英文摘要

Can a strictly local, iterated, weight-shared computation primitive support language modelling, and which of those three properties actually drives the model's behaviour? We define \textsc{TextNCA}, a 1D causal windowed-attention realisation of the Neural Cellular Automaton primitive, and study a hierarchical variant that cascades three stages with windows $w \in \{8, 32, 128\}$ and $T_s$ shared-weight iterations per stage, all on WikiText-103 at roughly 30M parameters and 60k training steps. The model does not match a parameter-matched Transformer at this scale (Hier-TextNCA $60.3$ vs.\ Transformer-6L $52.8$ and Transformer-12L $44.7$ PPL), so we treat it as an analytical probe rather than a proposed alternative. The behaviour we observe is largely explained by the staged narrow-to-wide schedule: a non-iterating sliding-window Transformer that reuses the same schedule comes within $+4.1$ PPL of the iterated model, while reversing, flattening, or breaking the monotonic ordering of the schedule costs between $+16.7$ and $+70.8$ PPL. Iteration adds a smaller bounded benefit on top of the schedule, with a clear optimum at $T_s{=}4$ and a U-shaped degradation beyond it. The GRU gate and learned per-step embeddings are required for that benefit to appear, and training with random $T_s$ yields an inference-time iteration-count knob at the cost of substantially higher absolute PPL. We position the work as a controlled reading of which parts of NCA-style computation carry the weight in language modelling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑