发表机构
Imperial College London(帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对带宽受限FPGA的流推理瓶颈,提出CascadeLUT框架,通过特征有序子集逐步预测减少数据移动,在延迟、吞吐量、能耗等指标上显著优于现有基线,还实现了低开销的设备端量化。
AI 中文摘要
将神经网络映射到FPGA可实现低延迟、高能效的推理,尤其适用于查找表(LUT)模型,这类模型无需乘法器,可直接映射到可重构结构。现有研究虽实现了高计算效率,但通常假设全样本可用,在带宽受限的流场景中会导致流水线停顿,此时瓶颈从计算转向数据移动,大量输入传输会限制吞吐量和能效。本文提出CascadeLUT,这是一种围绕带宽约束组织的信息结构化推理框架。该框架无需缓冲完整输入,而是将特征划分为有序子集,在子集到达时逐步细化预测;级联机制静态控制哪些层使用传入特征,实现无运行时分支的确定性流推理。通过将特征调度与硬件数据流协同设计,CascadeLUT在保持精度的同时减少了数据移动。在各数据集上,与现有LUT基线相比,它实现了4.0至12.5倍的延迟降低、3.0至5.0倍的吞吐量提升,以及最高13.8倍的每样本能耗降低,且每个任务仅使用最小DWN基线1.2至4.4倍的LUT资源。此外,本文还展示了集成LUT推理的设备端输入量化,并给出了真实 workload的端到端FPGA结果,量化开销降低了5倍。
英文摘要
Mapping neural networks to FPGAs enables low-latency, energy-efficient inference, particularly for lookup table (LUT)-based models that eliminate multipliers and map directly to reconfigurable fabric. While prior work achieves high compute efficiency, it typically assumes full-sample availability, causing pipeline stalls in bandwidth-limited streaming scenarios. Here, the bottleneck shifts from computation to data movement, as large input transfers limit throughput and energy efficiency. We present CascadeLUT, an information-structured inference framework organized around bandwidth constraints. Instead of buffering the full input, features are partitioned into ordered subsets and predictions are progressively refined as subsets arrive. The cascade statically controls which layers consume incoming features, enabling deterministic streaming inference without runtime branching. By co-designing feature scheduling with hardware dataflow, CascadeLUT reduces data movement while maintaining accuracy. Across datasets, it achieves 4.0 to 12.5 times lower latency, 3.0 to 5.0 times higher throughput and up to 13.8 times lower energy/sample than prior LUT baselines, using 1.2 to 4.4 times the LUTs of the smallest DWN baseline per task. We also demonstrate on-device input quantization integrated with LUT-based inference and present end-to-end FPGA results on real-world workloads, with 5 times reductions in quantization overhead.
CommentsTo appear in the proceedings of the 36th International Conference on Field-Programmable Logic and Applications (FPL 2026). 6 pages, 2 figures