arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11288cs.AR

基于概率内存计算硬件的仿生学习与决策:第二部分

Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 2

Thomas Dalgaty, Eiji Kawasaki, Miguel de Prado, Tommaso Salvatori, Germain Haugou, Eric Flamand

首次发表
浏览论文内容

中文总结 AI 辅助

本报告针对大规模概率能量模型在GPU上因HBM接口导致的延迟问题,提出采用概率模拟内存计算(AIMC)处理器,通过消除HBM接口实现超过1000倍的加速,验证了该架构的可扩展性。

中文摘要 AI 辅助

本报告扩展了我们之前的工作(第一部分),该工作介绍了一种用于不确定性下学习和决策的基于能量的模型。该模型利用随机朗之万动力学,持续演化神经元状态和模型权重上的近似概率分布。然而,正如第一部分所述并通过基于GPU的实现所证实,这种大规模概率能量模型由于过长的执行延迟而面临显著的可扩展性挑战。这种延迟源于一个根本性的不匹配:具有低算术强度的大规模并行模型(如基于能量的模型)在依赖高带宽内存(HBM)接口的处理器架构(如GPU)上执行。HBM对本质上可并行的模型施加了极其严格的顺序执行约束,造成此类模型不可扩展的错误印象。实际上,正是GPU架构本身,及其对HBM接口的依赖,对于这类AI模型而言并非可扩展的处理器架构。在本报告中,我们通过一个概率模拟内存计算(AIMC)处理器的详细事务级模型(TLM)证明,相同的基于能量的模型通过消除HBM接口并在片上内存中直接执行计算,其执行速度可以比数据中心级硬件快1000倍以上。

英文摘要

This report extends our previous work (Part 1), which introduced an energy-based model for learning and decision-making under uncertainty. The model leverages stochastic Langevin dynamics to continuously evolve approximate probability distributions over neuron states and model weights. However, as noted in Part 1 and confirmed through GPU-based implementations, large-scale probabilistic energy-based models of this nature face significant scalability challenges due to excessive execution latency. This latency stems from a fundamental mismatch: massively parallel models with low arithmetic intensity (such as energy-based models) are being executed on processor architectures like GPUs that rely on high-bandwidth memory (HBM) interfaces. The HBM imposes brutally sequential execution constraints on inherently parallelizable models, creating the false impression that such models are unscalable. In reality, it is the GPU architecture itself, with its dependence on HBM interfaces, that is not a scalable processor architecture for this class of AI model. In this report, we demonstrate using a detailed transaction-level model (TLM) of a probabilistic analogue in-memory computing (AIMC) processor that the same energy-based model can execute well over 1000x faster than data-center-grade hardware by eliminating the HBM interface and performing computation directly within on-chip memory.

↑