用于蒙特卡洛树搜索的多原语内存计算
Multi-primitive in-memory computing for Monte Carlo tree search
- Duke University(杜克大学)
- Hewlett Packard Labs(惠普实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对蒙特卡洛树搜索在传统处理器上能耗高限制边缘部署的问题,引入相到原语分解,将算法阶段映射为硬件原生IMC原语,应用于MCTS,实现高效能,在22纳米工艺下能效比CPU高96倍,比GPU高65至2059倍,同一基板可运行多领域应用。
AI中文摘要:
蒙特卡洛树搜索(MCTS)可实现人工智能决策,但在传统处理器上需要55-300瓦,限制了边缘部署。内存计算(IMC)在常规工作负载上节能,但被认为与不规则多阶段算法不兼容。我们引入了相到原语分解,将每个算法阶段重新表述为硬件原生IMC原语。应用于MCTS时,选择、扩展、模拟和反向传播分别对应内容可寻址内存、组合逻辑、电阻式随机存取存储器(RRAM)交叉开关和静态随机存取存储器,使搜索在芯片上进行。在22纳米工艺下,使用制造的RRAM阵列参数,IMC-MCTS在9x9围棋中功耗约为60毫瓦,相对于中央处理器(CPU)实现了96倍的能效提升,相对于H100图形处理器(GPU)实现了65倍至2059倍的能效提升。它在开源围棋引擎(Pachi-UCT和Michi-C)的样本大小不确定性范围内达到了欧洲围棋联盟的评级。同一基板可运行四个人工智能领域的八个应用程序。
英文摘要:
Monte Carlo tree search (MCTS) enables artificial intelligence (AI) decision-making, but requires 55-300 W on conventional processors, limiting edge deployment. In-memory computing (IMC) is energy-efficient on regular workloads but has been considered incompatible with irregular multi-phase algorithms. We introduce phase-to-primitive decomposition, which reformulates each algorithmic phase as a hardware-native IMC primitive. Applied to MCTS, selection, expansion, rollout and backpropagation map to content-addressable memory, combinational logic, a resistive random-access memory (RRAM) crossbar and static random-access memory, keeping search on chip. At 22 nm with fabricated RRAM-array parameters, IMC-MCTS consumes ~60 mW for 9x9 Go, achieving 96x energy efficiency over a central processing unit (CPU) and 65x-2,059x over an H100 graphics processing unit (GPU). It reaches a European Go Federation rating within sample-size uncertainty of open-source Go engines (Pachi-UCT and Michi-C). The same substrate runs eight applications across four AI domains.