arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

稀疏上限:脉冲网络可在何处及为何用活动换能量

The Sparsity Ceiling: Where Spiking Networks Can and Cannot Trade Activity for Energy

Zeyu Wang

arXiv 2607.26648首次发表:更新:

发表机构

Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究揭示脉冲网络稀疏上限由任务类型决定,前馈感知可低至5%放电率,循环语言模型约50%,脉冲Transformer可至2%,并提出信息论边界解释该现象,明确神经形态硬件在事件驱动感知的优势。

AI 中文摘要

脉冲神经网络(SNN)被视为高能效的载体,因为稀疏、事件驱动的活动将密集乘积累加替换为廉价的累加操作。我们认为,稀疏带来的能量增益并非SNN的固有属性,而是任务的特性。在保持架构固定的前提下,仅替换隐藏单元(连续型与漏极积分放电型),并采用双侧目标放电率探针,我们测量了活动可被压低至何种程度才会导致性能下降。低负载前馈感知网络可稀疏至5%的放电率且无精度损失;而循环语言模型无法低于约50%——循环状态必须保持活跃以承载信息。相比之下,脉冲Transformer可自由稀疏至2%(3个随机种子)——因此该上限是循环压缩的属性,而非序列建模的属性。注意力机制仅通过存储完整键值缓存摆脱了下限,用放电下限换取了内存开销:在神经形态硬件上,循环与注意力在不同维度产生开销,二者均无法摆脱。我们用信息论界限定式化该上限,公式为ρ ≥ H_b⁻¹(log₂M / H),并验证了其预测:下限随内存负载上升而升高,随状态宽度增加而降低,且(反驳了仅基于内存的朴素解读)随任务难度上升而升高。分层输入下限进一步限制了密集输入下的操作减少,将事件驱动感知隔离为神经形态硬件的优势领域。

英文摘要

Spiking neural networks (SNNs) are promoted as an energy-efficient substrate because sparse, event-driven activity replaces dense multiply-accumulates with cheap accumulates. We argue the energy dividend of sparsity is not a property of SNNs but of the task. Holding architecture fixed and swapping only the hidden unit (continuous vs. leaky-integrate-and-fire), plus a two-sided target-firing-rate probe, we measure how far activity can be pushed down before quality breaks. Low-load feed-forward perception sparsifies to 5% firing at no accuracy cost; a recurrent language model cannot go below ~50% -- the recurrent state must stay active to carry information. A spiking Transformer, by contrast, sparsifies freely to 2% (3 seeds) -- so the ceiling is a property of recurrent compression, not sequence modeling. Attention escapes the floor only by storing the full key-value cache, trading a firing floor for a memory wall: on neuromorphic hardware, recurrence and attention pay on different axes, neither escapes. We formalize the ceiling with an information-theoretic bound rho >= H_b^{-1}(log2 M / H) and confirm its predictions: the floor rises with memory load, falls with state width, and (refuting a naive memory-only reading) rises with task difficulty. A layer-wise input floor further caps op reduction under dense input, isolating event-driven perception as where neuromorphic hardware wins.

Comments5 pages, 6 figures. Code: https://github.com/zeyuyuyu/sparsity-ceiling

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑