arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Wiola 13M:一种用于参数高效小型语言模型的门控螺旋注意力架构

Wiola 13M, a Gated Spiral Attention Architecture for Parameter Efficient Small Language Models

Aryuemaan Kumar Chowdhury, Praveen Oosa, Vineesha Reddy

arXiv 2608.14604首次发表:更新:

AI 中文总结

针对10M-100M参数小型语言模型未适配小规模Transformer的问题,提出含三个创新组件的Wiola 13M模型,验证其等价性并发布开源实现,提升长程区分与梯度流性能。

AI 中文摘要

参数规模在1000万到1亿之间的小型语言模型,对于设备端推理、快速实验和可控科学研究颇具吸引力,但大多数此类模型都复用了标准Transformer块,未针对小参数规模进行适配。本文提出Wiola,这是一种仅解码器的语言模型,其创新点集中在每一层的三个即插即用组件上:第一,螺旋旋转位置编码,通过随维度缓慢增长的因子扰动标准旋转频率,使相位轨迹向外发散,在不增加参数的情况下提升长程区分能力;第二,门控螺旋注意力,引入基于查询流因果累积统计量的每头内容自适应标量门,以极低的代价提供隐式可微的软头选择形式;第三, Butterfly前馈块,用乘法交互和块内旁路路径替代传统扩展层,在匹配4倍门控线性单元块参数数量的同时,改善浅层堆叠中的梯度流。我们对每个组件进行了形式化定义,推导了精确的参数和计算预算,并证明门控注意力在全序列训练与缓存自回归解码之间存在经数值验证的精确等价关系,因此推理时不会引入任何近似。我们还描述了在标准微型故事语料库上的完全可复现训练与评估协议,参考实现作为开源包发布,权重已准备好用于推理支持。

英文摘要

Small language models in the ten to one hundred million parameter range are attractive for on device inference, rapid experimentation, and controlled scientific study, yet most of them reuse the standard transformer block without adaptation to the small scale regime. We present Wiola, a decoder only language model whose novelty is concentrated in three drop in components of every layer. First, Spiral Rotary Positional Encoding perturbs the standard rotary frequencies by a slowly growing per dimension factor so that phase trajectories fan outward, improving long range discrimination while adding no parameters. Second, Gated Spiral Attention introduces a per head, content adaptive scalar gate derived from a causal cumulative statistic of the query stream, providing an implicit and differentiable form of soft head selection at negligible cost. Third, the Butterfly feed forward block replaces the conventional expansion layer with a multiplicative interaction and an intra block bypass path, matching the parameter count of a four times gated linear unit block while improving gradient flow in shallow stacks. We formalize each component, derive exact parameter and computation budgets, and prove that the gated attention admits an exact and numerically verified equivalence between full sequence training and cached autoregressive decoding, so that no approximation is introduced at inference time. We also describe a fully reproducible training and evaluation protocol on a standard tiny story corpus. The reference implementation is released as an open source package with weights ready publishing support.

Comments6

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑