arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MAGMA:混合模型自适应高斯模型加速

MAGMA: Mixture-Model Adaptive Gaussian Model Acceleration

Peter Forcha, Harshitha Kajekusumadhar, Mbua Peter, Muhammed Kawser, Audrey Cyriell Mo

arXiv 2608.18366首次发表:更新:

AI 中文总结

针对传统FPGA-GMM加速器无法适应边缘系统场景变化的问题,提出MAGMA架构,实现GMM推理与在线EM参数自适应,在商用FPGA上达到显著加速效果并提升非平稳场景下的像素精度。

AI 中文摘要

传统基于FPGA的高斯混合模型(GMM)加速器使用离线训练的固定参数,限制了其在长寿命边缘系统中适应不断变化的场景统计特性的能力。我们提出MAGMA,一种完全可综合的定点FPGA架构,可从流式RGB像素输入执行并发GMM推理和在线期望最大化(EM)参数自适应。MAGMA将流水线推理数据路径与背景更新引擎相结合,采用硬件友好的超越近似方法——范围缩减切比雪夫指数、基于CLZ的对数运算以及移位减法除法器——同时配备防止方差崩溃和簇消亡的机制,以稳定在线定点EM。在AMD Spartan-7 XC7S50上实现,使用K=4个簇,MAGMA以74.49 MHz运行,占用7779个LUT、91个DSP,无块RAM,功耗274 mW。与软件相比,它实现了11.8倍的推理加速和81倍的M步加速,而空间子采样将每次更新的像素量减少了40倍,对EM收敛的影响极小。在合成非平稳场景下,MAGMA的在线自适应将平均像素精度优于静态基线(81.5% vs. 79.7%),证明在商用边缘FPGA上可实现完整的在线GMM学习。

英文摘要

Conventional FPGA-based Gaussian Mixture Model (GMM) accelerators use offline-trained, fixed parameters, limiting their ability to adapt to evolving scene statistics in long-lived edge systems. We present MAGMA, a fully synthesizable fixed-point FPGA architecture that performs concurrent GMM inference and online Expectation-Maximization (EM) parameter adaptation from a streaming RGB pixel input. MAGMA combines a pipelined inference datapath with a background update engine using hardware-friendly transcendental approximations---a range-reduced Chebyshev exponential, a CLZ-based logarithm, and a shift-and-subtract divider---alongside guards against variance collapse and cluster death that stabilize online fixed-point EM. Implemented on an AMD Spartan-7 XC7S50 with $K=4$ clusters, MAGMA runs at 74.49~MHz using 7,779 LUTs, 91 DSPs, and no block RAM, consuming 274~mW. It achieves an $11.8\times$ inference speedup and an $81\times$ M-step speedup over software, while spatial subsampling reduces per-update pixel volume by $40\times$ with minimal impact on EM convergence. Under a synthetic non-stationary scene, MAGMA's online adaptation improves mean pixel accuracy over a static baseline (81.5\% vs.\ 79.7\%), demonstrating that full online GMM learning is achievable on a commodity edge FPGA.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑