arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种面向硬件的高效贝叶斯推理计算与部署方法

A Hardware-oriented Approach for Efficient Bayesian Inference Computation and Deployment

Nikola Pižurica, Matteo Risso, Nikola Milović, Alessio Burrello, Igor Jovančević, Conor Heins, Miguel de Prado

arXiv 2607.17855首次发表:更新:

发表机构

PRAESC, Biel, Switzerland; Computer Science Center, University of Montenegro; Department of Control and Computer Engineering, Politecnico di Torino; Fain Tech, Podgorica, Montenegro(瑞士比耶尔预测分析与决策支持中心; 黑山大学计算机科学中心; 都灵理工大学控制与计算机工程系; 黑山波德戈里察法因科技公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对贝叶斯推理在边缘设备部署的计算成本问题,提出面向硬件的方法,通过重组内存布局、引入稀疏表示等优化算法,并结合自动调谐器,在NVIDIA Jetson Orin AGX上实现显著加速,输出与基线相同。

AI 中文摘要

贝叶斯推理为不确定性推理提供了原则性基础,但其计算成本阻碍了在资源受限的边缘设备上的部署。本文提出一种面向硬件的方法,用于加速商用现成嵌入式GPU上的离散贝叶斯推理。发现一类变分消息传递算法的延迟主要由张量收缩决定。通过两种互补合并策略重组操作的内存布局,引入可选的稀疏数组表示和张量聚类方案以减少内存占用。实例化该方法并为隐马尔可夫模型生成三种消息传递算法的优化变体,还辅以基于机器学习的自动调谐器。在NVIDIA Jetson Orin AGX上针对770个随机采样的现实部分可观测马尔可夫决策过程配置进行基准测试,实现了高达5倍的加速,典型增益为2 - 2.5倍,且输出与基线实现数值相同。

英文摘要

Bayesian inference provides a principled foundation for reasoning under uncertainty, but its computational cost hinders deployment on resource-constrained edge devices. In this paper, we present a hardware-oriented methodology for accelerating discrete Bayesian inference on commercial off-the-shelf embedded GPUs. We identify that the latency of a broad class of variational message-passing algorithms is dominated by tensor contractions. Our approach restructures the memory layout of these operations using two complementary merging strategies that produce compact, regularly-shaped primitives better suited for efficient GPU execution. We then introduce optional sparse array representations and a tensor-clustering scheme to reduce the memory footprint. We instantiate the methodology and produce optimized variants of three message-passing algorithms for Hidden Markov Models (HMMs), namely variational filtering, variational message passing, and marginal message passing. Furthermore, we complement this with a machine-learning-based autotuner that automatically selects the best-performing algorithmic variant for a given generative model specification. Benchmarked on an NVIDIA Jetson Orin AGX across 770 randomly sampled realistic Partially Observable Markov Decision Process (POMDP) configurations, our implementations achieve speedups of up to 5x, with typical gains of 2-2.5x, while producing numerically identical outputs to the baseline implementations.

CommentsCorrected the affiliation of Conor Heins. No changes to the scientific content

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑