arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20497cs.LGcs.AR

Bern2Edge:一种用于边缘部署的神经符号编译器,基于伯恩斯坦多项式网络

Bern2Edge: A Neurosymbolic Compiler for Edge Deployment via Bernstein Polynomial Networks

  • University of California, Irvine(加利福尼亚大学欧文分校)

机构由 AI 辅助整理,请以论文原文为准。

Malak Gamal El-Din, Yifan Zhang, Yasser Shoukry, Sitao Huang, Salma Elmalaki

AI总结:

本文提出Bern2Edge端到端框架,通过知识蒸馏将预训练教师网络转为硬件高效的伯恩斯坦多项式激活表示,支持两种部署路径,在边缘FPGA上实现了延迟、资源占用降低,同时保持较高准确率。

AI中文摘要:

在资源受限的边缘设备上部署高精度神经网络仍面临挑战,因为现有方法将训练、压缩和硬件综合视为独立阶段,在软件训练的模型与高效端到端部署之间存在差距,且对可解释性的支持有限。本文提出Bern2Edge,一种端到端框架,它利用知识蒸馏将预训练的教师前馈网络通过伯恩斯坦多项式激活函数转换为硬件高效的表示。该表示支持两种部署路径:(i)基于查找表(LUT)的高保真实现,在压缩下保留模型保真度;(ii)从伯恩斯坦激活几何中导出的基于符号规则的表示,支持具有显式输入空间约束的可解释推理。所得的BNN在相同压缩约束下,比ReLU实现了高达2.12个百分点(pp)的准确率提升。在系统层面,相对于AMD Xilinx KV260 FPGA上的W8A8量化教师,Bern2Edge实现了高达99.8%的延迟降低和95.2%的BRAM减少,同时准确率保持在0.5pp以内,还可进一步部署在低功耗Spartan-7 XC7S15 FPGA上。基于规则的路径将DSP使用量降低了高达89.0%,总准确率损失为1.5pp。

英文摘要:

Deploying high-accuracy neural networks on resource-constrained edge devices remains challenging, as existing approaches treat training, compression, and hardware synthesis as separate stages, leaving a gap between software-trained models and efficient end-to-end deployment with limited support for interpretability. We propose Bern2Edge, an end-to-end framework that uses knowledge distillation to convert a pretrained teacher feed-forward network into hardware-efficient representations via Bernstein polynomial activations. This representation enables two deployment paths: (i) a high-fidelity LUT-based realization that preserves model fidelity under compression, and (ii) a symbolic rule-based representation derived from Bernstein activation geometry, enabling interpretable inference with explicit input-space constraints. The resulting BNNs achieve up to 2.12 percentage-point (pp) accuracy improvement over ReLU under identical compression constraints. At the system level, Bern2Edge achieves up to 99.8% latency reduction and 95.2% BRAM reduction relative to a W8A8 quantized teacher on an AMD Xilinx KV260 FPGA, while maintaining accuracy within 0.5 pp, and further deploys on a low-power Spartan-7 XC7S15 FPGA. The rule-based path reduces DSP usage by up to 89.0% at a cost of 1.5 pp in total accuracy.

补充信息

↑