arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16500cs.LGcs.DCcs.DS

隐半马尔可夫模型的维特比算法的高性能张量表述

High-Performance Tensor Formulation of the Viterbi Algorithm for Hidden Semi-Markov Models

  • Sapienza University of Rome(罗马第一大学)

机构由 AI 辅助整理,请以论文原文为准。

Lorenzo Piarulli, Elia Belli, Daniele De Sensi

AI总结:

提出基于张量的HSMM维特比算法,将内循环映射到SIMD和并行架构,实现单核14倍、多核200倍、GPU570倍加速。

AI中文摘要:

隐半马尔可夫模型(HSMM)是基础的概率模型,在从计算生物学到金融和信号处理等多个领域中被广泛采用。维特比算法在给定HSMM的情况下解码最可能的状态序列,并可迭代应用于从头(ab initio)模型学习。然而,现有的维特比实现仍然是顺序执行的,且完全缺乏GPU加速的解决方案,这使得HSMM解码对于大规模工作负载而言不切实际。我们提出了一种基于张量的HSMM维特比算法表述,将内层循环重构为张量操作,这些操作自然地映射到SIMD单元和大规模并行架构上。基于这一表述,我们提供了涵盖单核和多核CPU以及首次在GPU上的优化实现。实验评估表明,与最先进的顺序基线相比,单核上加速比高达14倍,多核上超过200倍,GPU上超过570倍,为大规模HSMM解码建立了新的性能基线。

英文摘要:

Hidden Semi-Markov Models (HSMMs) are fundamental probabilistic models widely adopted across diverse domains, from computational biology to finance and signal processing. The Viterbi algorithm decodes the most likely state sequence given an HSMM and can be applied iteratively for ab initio model learning. However, existing Viterbi implementations remain sequential, and GPU-accelerated solutions are entirely absent, making HSMM decoding impractical for large-scale workloads. We present a tensor-based formulation of the Viterbi algorithm for HSMMs, restructuring the inner loops into tensor operations that naturally map onto SIMD units and massively parallel architectures. Building on this formulation, we provide optimized implementations spanning single- and multi-core CPUs, and, for the first time, GPU. Experimental evaluation demonstrates speedups of up to 14x on a single core, over 200x with multi-core, and over 570x on GPU over the state-of-the-art sequential baseline, establishing a new performance baseline for large-scale HSMM decoding.

↑