AI 中文总结
TriPLU替换小型语言模型的门控前馈网络分支为3阶乘积分支,在低计算场景下提升了TinyStories等数据集的验证损失,但对优化敏感且未确立通用效率。
AI 中文摘要
我们研究仅含解码器的小型语言模型是否能从直接相乘学习到的特征投影的前馈层中获益。TriPLU(三线性乘积线性单元)将常用的门控前馈网络分支替换为仅乘积的3阶分支,该分支按坐标对三个投影流进行相乘。在字符级TinyStories 1兆字节前缀研究中,TriPLU达到的平均最佳验证损失为1.0637,而匹配度相近的SwiGLU为1.1017,4阶乘积对照组为1.0780,2阶对照组为1.1026。在仅训练的Byte-BPE实验中,TriPLU在低学习率设置下还降低了TinyStories和WikiText-2原始数据的验证及保留字节每比特数,PMI切片证据表明其在已见过的中高PMI相邻词对上有提升。恒定学习率诊断显示,乘积分支归一化可降低高学习率下的最佳检查点差距,不过在热调度下最终字节每比特数仍会下降。所得结论刻意限定范围:直接乘积前馈网络可在特定低计算场景下提升固定预算小型模型的损失,但该分支对优化敏感,未确立浮点运算归一化效率、缩放行为或广泛的大语言模型性能。
英文摘要
We study whether tiny decoder-only language models benefit from feed-forward layers that directly multiply learned feature projections. TriPLU, a Trilinear Product Linear Unit, replaces the usual gated FFN branch with a product-only degree-3 branch that multiplies three projected streams coordinatewise. In a character-level TinyStories 1M-byte prefix study, TriPLU reaches a mean best validation loss of 1.0637, compared with 1.1017 for closely matched SwiGLU, 1.0780 for a degree-4 product control, and 1.1026 for a degree-2 control. In train-only Byte-BPE experiments, TriPLU also lowers validation and heldout bits per byte on TinyStories and WikiText-2 raw under low-learning-rate settings, with PMI-slice evidence suggesting gains on seen middle- and high-PMI adjacent-token pairs. Constant-learning-rate diagnostics show that product-branch normalization can reduce the high-learning-rate best-checkpoint gap, although final BPB still degrades under hot schedules. The resulting claim is deliberately narrow: direct product FFNs can improve fixed-budget small-model loss in specific low-compute regimes, but the branch is optimization-sensitive and does not establish FLOP-normalized efficiency, scaling behavior, or broad LLM performance.