arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03656cs.SDcs.AI

重新审视人声合奏多音高估计中的输入时频表示

Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles

  • Yonsei University(延世大学)
  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

Junyoung Koh, Hao-Wen Dong

AI总结:

本文重新审视人声合奏多音高估计中的输入时频表示,发现线性STFT在性能和效率上优于HCQT,且更精细的频率分辨率并非必要,较短分析窗口更有效。

AI中文摘要:

在人声合奏中,多音高估计具有挑战性,因为歌手占据重叠的音高范围,且常常以间隔很近的基频演唱,导致其谐波在时频表示中重叠。现有模型通常使用基于谐波常数Q变换(HCQT)的表示来提供频率自适应分辨率,但当训练混合数据实时生成时,这会以昂贵的特征提取为代价。我们重新审视了这一设计,并将HCQT与线性短时傅里叶变换(STFT)进行比较,后者将频率 bins 直接作为模型输入。尽管线性STFT具有固定的频率分辨率且缺乏音高对齐的输入网格,但其性能优于HCQT,同时大幅降低了特征提取成本。进一步分析表明,更长的分析窗口或更宽的频谱覆盖范围并未带来额外改进,而将输入限制在预测音高范围内则削弱了线性STFT的优势。这些结果表明,更精细的频率分辨率不一定能改善人声合奏的MPE,且较短的分析窗口可能对时变的人声音高更有效。

英文摘要:

Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency representations. Existing models commonly use harmonic constant-Q transform (HCQT)-based representations to provide frequency-adaptive resolution, at the cost of expensive feature extraction when training mixtures are generated on the fly. We revisit this design and compare HCQT with a linear short-time Fourier transform (STFT), whose frequency bins are directly provided as model inputs. Despite its fixed frequency resolution and the absence of a pitch-aligned input grid, the linear STFT outperforms HCQT while substantially reducing feature-extraction cost. Further analysis shows that a longer analysis window or broader spectral coverage provides no additional improvement, while restricting the input to the predicted pitch range reduces the advantage of the linear STFT. These results suggest that finer frequency resolution does not necessarily improve vocal-ensemble MPE, and that shorter analysis windows can be more effective for time-varying vocal pitches.

↑