arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29119eess.SPcs.NAcs.SDeess.ASmath.NA

可逆常数Q变换的精确分解与快速计算

Exact Factorisation and Fast Computation of Invertible Constant-Q Transforms

Facundo Franchino, Eloi Moliner, Vesa Välimäki

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Flash-CQT,通过精确分解实现可逆常数Q变换的快速GPU计算,显著降低计算时间和内存占用,适用于音频分析与机器学习。

中文摘要 AI 辅助

常数Q变换(CQT)在对数频率轴上表示音频。其非平稳Gabor形式是精确可逆的,但各频带中时间系数数量不等,使GPU计算复杂化。一种精确分解将频谱选择、共轭、加窗和重排组合成一个固定映射,连接一个打包傅里叶变换与较短的频带逆变换。这些因子提供了波形重建、用于反向传播的实伴随算子,以及算术深度和块宽度的界限;重叠切片支持有界内存的流式处理。在两个GPU型号上的测试表明,Flash-CQT将分析-合成往返时间相对于计算相同CQT的基线减少了2到8倍。所提出的实现还使用了超过30%更少的峰值临时工作空间,并在单精度浮点运算中达到了可忽略的重建误差,信噪比约为130 dB。这些进展使Flash-CQT成为频谱分析和现代音频机器学习系统中实用且计算高效的前端。

英文摘要

The constant-Q transform (CQT) represents audio on a logarithmic frequency axis. Its nonstationary Gabor formulation is exactly invertible, but the unequal numbers of time coefficients in its bands complicate GPU computation. An exact factorisation combines spectral selection, conjugation, windowing, and reordering into a fixed map between one packed Fourier transform and the shorter band inverse transforms. The factors give waveform reconstruction, real adjoints for backpropagation, and bounds on arithmetic depth and block width; overlapping slices permit streaming with bounded memory. Tests on two GPU models show that Flash-CQT reduces analysis-synthesis round-trip time by factors of two to eight relative to a baseline computing the same CQT. The proposed implementation also uses over 30% less peak temporary workspace and reaches a negligible reconstruction error, with a signal-to-noise ratio of about 130 dB, in single-precision floating-point arithmetic. These advances make Flash-CQT a practical, computationally efficient front end for spectral analysis and modern audio machine-learning systems.

发表机构

  • Massachusetts Institute of Technology(麻省理工学院)
  • Aalto University(阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑