arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CircuitsDNA:通过进化合成发现非常规多精度算术电路

CircuitsDNA: Discovering Unconventional Multi-Accuracy Arithmetic Circuits via Evolutionary Synthesis

Ruichen Qi, Junyi Luo, Xinting Jiang, Quan Cheng, Gregory Kielian, Ben Laurie, Dennis Sylvester, Mehdi Saligane

arXiv 2609.01735首次发表:更新:

发表机构

Brown University; Google; University of Michigan(布朗大学; 谷歌; 密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CircuitsDNA是一种进化框架,可自动进化出单电路内多精度模式的算术电路,其8位乘法器在面积功耗比、收敛速度等方面优于传统方法,精度损失可控。

AI 中文摘要

新兴边缘AI workload日益需要可按需在计算精度与效率间权衡的算术单元。然而,现有近似算术电路通常是固定精度的,或依赖预定义结构实现运行时可配置性。本研究提出CircuitsDNA,一种进化框架,可自动进化出支持单电路内多种精度模式的精度可配置算术电路。它整合了三大核心特性:1)多阈值可验证斜接器,用于强制特定模式的精度要求;2)资源受限的可验证驱动搜索,在不牺牲正确性的前提下降低验证开销,实现对大型电路设计的高效探索;3)反馈驱动的自适应变异,用于优先考虑有效的结构修改并加速搜索收敛。实验结果显示,在28nm CMOS工艺下合成的8位乘法器变体,与精确8位乘法器相比,在INT8 DNN workload下面积-功耗乘积最多降低56%,在穷尽活动下最多降低93%。在 worst-case error(WCE)预算最多为1%的情况下,经过微调后,其相对于FP32的精度损失在CNN和DeiTs上均低于2%。CircuitsDNA消除了传统方法在8/12/16位乘法器上观察到的所有搜索停滞,且自适应变异相比非自适应对应物提供了最多1.33倍的收敛速度提升。

英文摘要

Emerging edge AI workloads increasingly require arithmetic units that can trade computational accuracy for efficiency on demand. However, existing approximate arithmetic circuits are typically fixed-accuracy or rely on predefined structures for runtime configurability. This work introduces CircuitsDNA, an evolutionary framework that automatically evolves accuracy-configurable arithmetic circuits supporting multiple accuracy modes within a single circuit. It integrates three key features: 1) multi-threshold verifiability miter to enforce mode-specific accuracy requirements, 2) resource-limited verifiability-driven search to reduce verification overhead without sacrificing correctness, enabling efficient exploration of large circuit design, and 3) feedback-driven adaptive mutation to prioritize effective structural modifications and accelerate search convergence. Experimental results show that the 8-bit multiplier variants synthesized in 28-nm CMOS reduce the area-power product by up to 56% on INT8 DNN workload and 93% under exhaustive activity, compared with an exact 8-bit multiplier. Across CNNs and DeiTs, the accuracy loss relative to FP32 remains below 2% after fine-tuning under worst-case error (WCE) budgets of at most 1%. CircuitsDNA eliminates all search stalls observed in conventional methods across 8/12/16-bit multipliers, while adaptive mutation provides up to 1.33 times faster convergence than its non-adaptive counterpart.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑