arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03058eess.AScs.SD

无监督瞬时相位与频率跟踪:基于逆语音合成方法

Unsupervised Instantaneous Phase and Frequency Tracking by Inverse Voice Synthesis

Chin-Yun Yu, György Fazekas

首次发表
浏览论文内容

中文总结 AI 辅助

针对神经声码器难以端到端学习基频的问题,提出无混叠加性声源的源-滤波器模型,显式建模瞬时相位并求导得瞬时频率,无需外部跟踪器即可实现高精度无监督基频跟踪,在多个数据集上优于神经基线。

中文摘要 AI 辅助

知识驱动的神经声码器难以端到端地学习可靠的基频,因为频谱目标对周期结构的监督较弱,且缺乏相位信息。我们通过一个源-滤波器模型来解决这一问题,该模型的无混叠加性声源使声门周期的瞬时相位显式化;对其求导即可得到瞬时频率,从而无需外部跟踪器即可获得基频($F_0$)。波形误差仅监督确定性谐波路径,而频谱损失则覆盖整个信号。在M4Singer和LM-SSD数据集上,重建结果实现了相位对齐,信号重建误差比达到8.1 dB,而神经基线方法仍为负值。然而,GOLF在给定外部$F_0$时,仍能达到更低的频谱失真。在LM-SSD上,恢复的$F_0$在所有测试方法中取得了最高的整体准确率,包括直接使用的监督神经音高跟踪器,且声门闭合时刻的识别率与REAPER相差仅0.53个百分点,而无需任何$F_0$标签。

英文摘要

Knowledge-driven neural vocoders struggle to learn reliable fundamental frequency end-to-end, because spectral objectives provide weak supervision of periodic structure and lack phase information. We address this with a source-filter model whose alias-free additive source makes the instantaneous phase of the glottal cycle explicit; differentiating it yields the instantaneous frequency, and thus $F_0$, without an external tracker. Waveform error supervises only the deterministic harmonic path, while a spectral loss covers the full signal. On M4Singer and LM-SSD, the reconstruction is phase-aligned, reaching a signal-to-reconstruction-error ratio of 8.1 dB, while neural baselines remain negative. However, GOLF, given an external $F_0$, still reaches lower spectral distortion. On LM-SSD, the recovered $F_0$ attains the highest overall accuracy of any method tested, including supervised neural pitch trackers applied off the shelf, and the glottal closure instants come within 0.53 points of REAPER's identification rate, without any $F_0$ label.

发表机构

  • Queen Mary University of London(伦敦玛丽女王大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑