arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FINNAS:面向FPGA喷注子结构分类的FINN引导硬件感知神经架构搜索与剪枝

FINNAS: FINN-Guided Hardware-Aware NAS and Pruning for FPGA Jet Substructure Classification

Eva Chauffour, Changhong Li, Georgios Floros, Shreejith Shanker

arXiv 2609.16367首次发表:更新:

发表机构

Reconfigurable Computing Systems Lab, Electronic & Electrical Engineering; Trinity College Dublin, Ireland(可重构计算系统实验室,电子与电气工程系; 都柏林三一学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FINNAS提出FINN引导的硬件感知进化神经架构搜索框架,联合搜索量化MLP深度、宽度和精度,在CERNBox喷注子结构分类任务上,相比手动优化FINN加速器,精度提升至74.36%,LUT减少8.5倍,延迟减少1.77倍。

AI 中文摘要

FPGA非常适合在严格的精度、延迟和资源约束下部署量化神经网络(QNN);然而,识别高效的模型-加速器组合通常需要大量的人工设计空间探索和重复的硬件综合。本文提出了FINNAS,一个FINN引导的硬件感知进化神经架构搜索框架。FINNAS联合搜索量化MLP的深度、宽度和全局精度设置,并在完全并行映射下,使用代理验证精度以及FINN估计的LUT使用量和延迟对候选方案进行排名。选定的最终候选方案被完全重新训练,进行搜索后的非结构化剪枝,并使用RTL仿真和Vivado上下文外综合进行验证。在CERNBox喷注子结构分类任务上,搜索到的实现展示了有竞争力的精度-资源权衡。与手动优化的密集FINN加速器相比,紧凑的FINNAS设计将精度从73.78%提高到74.36%,同时将LUT使用量减少8.5倍,RTL仿真延迟减少1.77倍。非结构化剪枝进一步在完全并行的最终候选方案中提供了一致的LUT和FF减少。

英文摘要

FPGAs are well suited to deploying quantised neural networks (QNNs) under strict accuracy, latency, and resource constraints; however, identifying efficient model-accelerator combinations commonly requires extensive manual design-space exploration and repeated hardware synthesis. This paper presents FINNAS, a FINN-guided hardware-aware evolutionary neural architecture search framework. FINNAS jointly searches quantised MLP depth, width, and global precision settings, and ranks candidates using proxy validation accuracy together with FINN-estimated LUT usage and latency under a fully parallel mapping. Selected finalists are fully retrained, subjected to post-search unstructured pruning, and validated using RTL simulation and Vivado out-of-context synthesis. On the CERNBox jet substructure classification task, the searched implementations expose competitive accuracy-resource trade-offs. Compared with a manually optimised dense FINN accelerator, a compact FINNAS design improves accuracy from 73.78\% to 74.36\%, while reducing LUT usage by \(8.5\times\) and RTL-simulation latency by \(1.77\times\). Unstructured pruning further provides consistent LUT and FF reductions across the fully parallel finalists.

CommentsAccepted by ICECS'26

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑