发表机构
SymbolicLight Research(SymbolicLight 研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SymbolicLight V2通过混合神经形态架构和稀疏执行,在FPGA和ARM上实现低能耗语言推理,显著降低解码能耗并提升吞吐量。
AI 中文摘要
SymbolicLight V2将稀疏事件计算与连续状态处理相结合,构建了一种混合神经形态语言架构。在V1的脉冲门控双路径基础上,它进一步在投影层添加了分级有符号事件,并引入了无softmax的局部注意力。我们使用数字定点算术在Alveo U50C FPGA上实现了这一1.94亿参数的模型,并在ARM CPU上通过稀疏整数执行进行了实现。在175 MHz频率下,三种相同检查点的FPGA实现中,主动行权重收集和有效状态KV加载将32个令牌前缀和128个输出的解码吞吐量从474.6提升至643.2令牌/秒。估算的整卡能耗从每个生成令牌0.06087焦耳降至0.04407焦耳,降幅达27.6%。包括预填充在内的完整请求能耗在三种前缀长度下降低了24.4%至27.7%。一项独立的空闲功耗分解显示,整卡能耗的82.8%归因于加载后的空闲状态,这解释了缩短令牌延迟带来的收益。与记录的RTX 5090编译FP32基线相比,整数FPGA执行在短上下文解码期间估算整卡能耗降低了89.1%;但算术精度存在差异,且GPU基线并非测试中能耗最低的配置。在ROCK 5T的四个Cortex-A76核心上,完整请求在适配器交流输入下达到65.4令牌/秒,功耗9.80瓦,每个生成令牌能耗0.151焦耳。这些结果将事件稀疏性与省略的计算和数据移动联系起来。相关机制也支持其他专用V2实现:吞吐量提升幅度大于有功功率增幅,从而降低每个生成令牌的能耗。评估保持部署检查点不变;其质量落后于同等预算的密集对照模型,因此结果并未证明在同等质量下的能效优势。
英文摘要
SymbolicLight V2 combines sparse event computation with continuous-state processing in a hybrid neuromorphic language architecture. Extending V1's spike-gated dual paths, it adds graded signed events at further projections and softmax-free local attention. We implement the 194M-parameter model on an Alveo U50C FPGA using digital fixed-point arithmetic and on an ARM CPU using sparse integer execution. Across three same-checkpoint FPGA implementations at 175 MHz, active-row weight gathering and valid-state KV loading raise decode throughput from 474.6 to 643.2 tokens/s for a 32-token prefix and 128 outputs. Estimated gross card energy falls from 0.06087 to 0.04407 J per generated token, a 27.6% reduction. Complete-request energy, including prefill, falls by 24.4-27.7% across three prefix lengths. An independent idle split attributes 82.8% of gross card energy to loaded idle, explaining the benefit of shorter token latency. Against the recorded RTX 5090 compiled-FP32 baseline, integer FPGA execution uses 89.1% less estimated card energy during short-context decode; arithmetic precisions differ, and the GPU baseline is not the lowest-energy tested configuration. On four Cortex-A76 cores of a ROCK 5T, complete requests reach 65.4 tokens/s at 9.80 W and 0.151 J per generated token at the adapter's AC input. These results connect event sparsity to omitted computation and data movement. The mechanisms also support other dedicated V2 implementations: increasing throughput by a greater factor than active power lowers energy per generated token. Evaluation holds the deployed checkpoint fixed; its quality trails a same-budget dense control, so the results do not establish equal-quality efficiency.
Comments22 pages, 8 figures, 11 tables