arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将强制对齐扩展到终端设备

Scaling Forced Alignment to End-User Devices

Lawry Sorenson, Michael Crandall, Eric K. Ringger, Stephen D. Richardson

arXiv 2609.21145首次发表:更新:

发表机构

Brigham Young University(杨百翰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出两种优化(Hirschberg算法和受约束随机游走剪枝),将强制对齐扩展到终端设备,大幅降低内存并提速,同时保持高准确性。

AI 中文摘要

维特比算法此前已被用于执行音频到文本的强制对齐,以从在线资源中挖掘训练数据。然而,许多现有实现具有二次方的时间和空间复杂度,难以扩展到长输入序列。我们提出两种优化来解决此问题。首先,我们应用Hirschberg算法以线性内存就地执行对齐。其次,我们将语音与文本之间的对齐建模为受约束的随机游走,从而能够在考虑转录错误的同时,以任意置信度剪枝搜索空间。Hirschberg优化将三小时输入的显存使用从140 GB降至5 MB,同时在CPU上运行时的对齐结果与torchaudio完全一致,且耗时仅为后者的三分之一。对于超过20分钟的输入,我们通过剪枝额外获得2倍加速,同时在超过98%的测试案例中保持对齐准确性。

英文摘要

The Viterbi algorithm has been previously used to perform forced alignment of audio to text to mine training data from online resources. However, many existing implementations have quadratic time and space complexity, scaling poorly to long input sequences. We propose two optimizations to address this issue. First, we apply the Hirschberg algorithm to perform the alignment in place using linear memory. Second, we model the alignment between speech and text as a constrained random walk, allowing us to prune the search space with arbitrary confidence while accounting for transcription errors. The Hirschberg optimization reduces memory usage from 140 GB to 5 MB for three-hour inputs while producing identical alignments in one-third the time of torchaudio when both run on a CPU. We achieve an additional 2x speedup with pruning on inputs longer than 20 minutes while preserving alignment accuracy in more than 98% of tested cases.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑