发表机构
Research Institute of Electrical Communication, Tohoku University; Graduate School of Engineering, Tohoku University(东北大学电气通信研究所; 东北大学研究生院工学研究科)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于差分随机模拟退火的2048自旋全连接处理器,通过仅重算翻转自旋交互和稀疏调度加速,在28nm工艺下实现2.7ms求解时间和0.86mJ能耗,功耗和能耗分别降低1.5倍和3.5倍。
AI 中文摘要
本文提出了一种基于差分随机模拟退火(DSSA)的2048自旋全连接退火处理器,该处理器采用台积电28纳米CMOS工艺进行架构设计,后布局面积为3毫米×4毫米。处理器在500 MHz频率下满足时序要求,集成了16 Mb SRAM权重存储器,并通过面积高效的随机数生成器将随机噪声分摊到16个自旋上。DSSA保持了串行数据通路以实现高密度,但仅对发生翻转的自旋重新计算相互作用,从而在每次退火运行中将有效工作量缩减至活跃前沿。自旋选择调度、基于优先级的权重读取以及跳过空闲步长的温度控制器加速了稀疏更新,同时不牺牲全连接性。后布局仿真结果显示,在2000自旋问题上,求解时间(TTS)为2.7毫秒,求解能耗为0.86毫焦耳,功耗为316毫瓦(每自旋0.15毫瓦),与先前预计的全连接退火器相比,功耗降低1.5倍,TTS能耗降低3.5倍。这些结果证明了所提出的DSSA架构在后布局评估下用于大规模组合优化硬件的潜力。
英文摘要
A 2,048-spin fully connected annealing processor based on differential stochastic simulated annealing (DSSA) is presented as an architectural design in TSMC 28 nm CMOS with a 3 mm x 4 mm post-layout area. The processor closes timing at 500 MHz, integrates a 16 Mb SRAM weight memory, and amortizes stochastic noise across 16 spins with area-efficient random number generators. DSSA keeps a serialized datapath for density but recomputes interactions only for spins that flip, shrinking the effective workload to the active frontier during each annealing run. Spin-select scheduling, priority-based weight reads, and a temperature controller that skips idle steps accelerate sparse updates without sacrificing full connectivity. Post-layout simulation results show 2.7 ms time-to-solution (TTS) and 0.86 mJ energy-to- solution on 2,000-spin problems at 316 mW (0.15 mW/spin), achieving 1.5x lower power and 3.5x lower TTS energy than projected prior fully connected annealers. These results demonstrate the potential of the proposed DSSA architecture for large-scale combinatorial optimization hardware under post-layout evaluation.
DOI:10.1109/ACCESS.2026.3731035