arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BitFair:一款具有可学习早期终止和自适应比特排序功能的12纳米比特串行CNN加速器,用于超低功耗XR视觉

BitFair: A 12-nm Bit-Serial CNN Accelerator with Learnable Early Termination and Adaptive Bit Ordering for Ultra-Low-Power XR Vision

Ang Li, Chang Gao

arXiv 2607.05445首次发表:更新:

发表机构

Delft University of Technology(代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对XR可穿戴设备低功耗、低延迟需求,提出BitFair比特串行CNN加速器,通过可学习早期终止和自适应比特排序,利用动态比特级稀疏性,实验表明其在准确率和能效上优于现有加速器。

AI 中文摘要

扩展现实(XR)可穿戴设备需要在几瓦的严格功率范围内实现始终在线感知,并且运动到光子的延迟预算低于20毫秒,留给神经网络推理的时间只有几毫秒。比特串行计算对于这种节能的神经网络加速很有吸引力,但许多现有架构即使在ReLU将最终输出设置为零时仍会处理所有比特。本文提出了BitFair,这是一款软硬件协同设计的比特串行CNN加速器,具有可学习的比特级早期终止和自适应比特排序功能,可在XR应用的超低功耗和严格延迟要求下工作。BitFair通过学习每层阈值来利用动态比特级稀疏性,当部分和可靠地预测最终ReLU输出将为零时触发早期终止。此外,它会搜索逐层的比特顺序,优先考虑信息丰富的比特,在不牺牲准确性的情况下最大化早期终止。采用GlobalFoundries 12nm FinFET工艺实现,核心面积为0.34平方毫米,片上内存为104KB,电压从0.55V扩展到0.70V,实现了亚毫秒级延迟、高达117.0 BTOPS/W的功率效率和0.07 pJ/SOP的能量消耗。在IBM DVS128手势和N-MNIST数据集上,BitFair分别实现了96.5%和97.7%的准确率,同时与之前制造的XR视觉加速器相比,有效能量效率提高了4.0-22.1倍,准确率提高了高达9.2%。

英文摘要

Extended Reality (XR) wearables require always-on perception within tight power envelopes of a few watts and motion-to-photon latency budgets below 20 ms, leaving only a few milliseconds for neural-network inference. Bit-serial computing is attractive for such energy-efficient neural network acceleration, but many existing architectures still process all bits even when ReLU sets the final output to zero. This paper presents BitFair, a software-hardware co-designed bit-serial CNN accelerator with learnable bit-level early termination and adaptive bit ordering, working under the ultra-low-power and strict latency requirements of XR applications. BitFair exploits dynamic bit-level sparsity by learning per-layer thresholds that trigger early termination when partial sums reliably predict that the final ReLU output will be zero. Furthermore, it searches for layer-wise bit orders that prioritize informative bits, maximizing early termination without sacrificing accuracy. A GlobalFoundries 12-nm FinFET implementation with a core area of 0.34 mm^2, 104 KB on-chip memory, and voltage scaling from 0.55 to 0.70 V achieves sub-millisecond latency, up to 117.0 BTOPS/W, and 0.07 pJ/SOP. On IBM DVS128 Gesture and N-MNIST, BitFair achieves 96.5% and 97.7% accuracy, respectively, while improving effective energy efficiency by 4.0-22.1x and accuracy by up to 9.2% over prior fabricated XR vision accelerators.

CommentsAccepted to IEEE Journal on Emerging and Selected Topics in Circuits and Systems

DOI:10.1109/JETCAS.2026.3722396

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑