发表机构
Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对脉冲语言模型的时间编码与非线性计算权衡问题,提出联合设计脉冲编码和注意力算子的Spora模型,在GLUE、CoLA基准上较SpikeLM取得显著性能提升,且编码保真度与计算成本的关联得到进一步表征。
AI 中文摘要
脉冲语言模型面临着在短时间窗口内表示连续语义特征与保留高成本非线性注意力操作之间的权衡问题。我们引入Spora,它联合设计了脉冲编码和注意力算子。二元时间权重使T个脉冲能够表示具有多达T位容量的组合值,相比之下,脉冲计数读出仅能提供O(log₂T)位。单极二元脉冲(UBS)使用阈值和脉冲触发的残差衰减来生成非负整数编码;双极二元脉冲(BBS)将符号与幅值分离,并为带符号激活学习一个尺度。这些表示支持注意力中的累加-移位点积和整数指数映射。在四个时间步长下,Spora实现了76.6的GLUE平均得分和44.1的CoLA MCC,分别比SpikeLM提高了1.2和6.2个百分点。将BBS扩展到六个步长时,这些得分提升至78.2和47.4。条件衰减分析、匹配预算的激活量化比较、事件工作负载统计以及定点评估进一步表征了编码保真度与计算成本之间的联系。
英文摘要
Spiking language models face a tradeoff between representing continuous semantic features over short temporal windows and retaining costly nonlinear attention operations. We introduce Spora, which jointly designs spike encodings and attention operators. Binary temporal weights let $T$ spikes represent compositional values with up to $T$ bits of capacity, compared with $O(\log_2 T)$ bits for spike-count readout. Unipolar Binary Spiking (UBS) uses thresholds and spike-triggered residual decay to produce non-negative integer codes; Bipolar Binary Spiking (BBS) separates sign and magnitude and learns a scale for signed activations. These representations support accumulation-and-shift dot products and integer-exponent mappings in attention. With four time steps, Spora achieves 76.6 average GLUE score and 44.1 CoLA MCC, improving over SpikeLM by 1.2 and 6.2 points, respectively. Extending BBS to six steps raises these scores to 78.2 and 47.4. Conditional-decay analysis, matched-budget activation-quantization comparisons, event-workload statistics, and fixed-point evaluation further characterize the connection between encoding fidelity and computational cost.