基于混沌可重构无时钟芯片的大规模并行强化学习
Massively Parallel Reinforcement Learning with a Chaotic Reconfigurable Clockless Chip
浏览论文内容
中文总结 AI 辅助
本文提出基于异步布尔网络的无时钟芯片架构,利用分布式混沌产生并行熵源,在1024臂老虎机问题上实现并行决策,并扩展至5120通道,生成速率达2.14 TS/s,为大规模强化学习提供高能效硬件基底。
中文摘要 AI 辅助
基于物理动态系统的硬件加速器为节能强化学习应用提供了一条有吸引力的途径。然而,其可扩展性具有挑战性,因为这需要许多统计上独立的熵源。在此,我们介绍一种基于异步布尔网络(或格)的准模拟决策架构,该架构在无时钟可重构芯片上实现。网络中的每个节点由单个逻辑元件组成,充当自主熵源。该架构产生分布式布尔混沌,其中空间耦合网络生成并行混沌布尔转换流,节点之间的统计依赖性非常低。我们实验性地在1024臂老虎机问题上展示了并行决策,这超出了先前硬件实现的范围,同时显著改善了幂律缩放性能。另外,我们将所提出的熵源扩展到5120个并行通道,产生2.14 TS/s的总样本生成速率。我们的解决方案在商业可重构CMOS芯片上实现,并提供高集成密度和易于编程性。我们的结果为使用分布式布尔混沌作为大规模强化学习的有价值硬件基底以及开发完全集成、高吞吐量决策加速器铺平了道路。
英文摘要
Hardware accelerators based on physical dynamical systems offer an attractive route toward energy-efficient reinforcement learning applications. However, their scalability is challenging because it requires many statistically independent entropy sources. Here, we introduce a quasi-analog decision-making architecture based on asynchronous Boolean networks (or lattices) implemented on a clockless reconfigurable chip. Each node in the network consists of a single logic element that acts as an autonomous entropy source. This architecture gives rise to distributed Boolean chaos, in which a spatially coupled network generates parallel streams of chaotic Boolean transitions with very low statistical dependence between nodes. We experimentally demonstrate parallel decision-making on a 1024-armed bandit problem, which is beyond the scale of previous hardware implementations, while significantly improving power-law scaling performance. Separately, we scale the proposed entropy source to 5120 parallel channels, yielding an aggregate sample generation rate of 2.14 TS/s. Our solution is implemented on a commercial reconfigurable CMOS chip and offers high integration density and ease of programmability. Our results pave the way for using distributed Boolean chaos as a valuable hardware substrate for large-scale reinforcement learning and for the development of fully integrated, high-throughput decision-making accelerators.
发表机构
- CentraleSupélec and Université de Lorraine(中央苏佩莱克学院和洛林大学)
机构由 AI 辅助整理,请以论文原文为准。