arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新审视Vocos:时频神经声码器中的相位问题

Revisiting Vocos: That Phasiness Business in Time-Frequency Neural Vocoding

Ünal Ege Gaznepoğlu, Frank Zalkow, Mohammad Joshaghani, Emanuël A. P. Habets, Nils Peters, Christian Dittmar

arXiv 2607.24323首次发表:更新:

AI 中文总结

研究从相位重建角度重新审视Vocos,用带限梅尔频谱图量化其时域与时频域声码器差距,经消融研究发现其对幅度建模有效对相位较差,调整主干预测相位差时一维卷积层有阻碍,指出未来研究应关注归纳偏差以更好建模语音信号时频结构。

AI 中文摘要

最近,时频神经声码器已接近时域神经声码器的先进质量。Vocos因其效率而成为一个显著例子,但其音频质量落后于时域声码器,原因仍存在争议。因此,在本研究中,我们从相位重建的角度重新审视Vocos。首先,我们使用带限梅尔频谱图作为输入来量化时域和声频域声码器之间的差距。随后,通过消融研究,我们验证了Vocos架构对幅度建模有效,但对相位建模效果较差。然后,我们调整Vocos主干以预测相位差(相位重建的前兆),并确定一维卷积层阻碍了其准确预测。我们的发现表明,未来的研究需要关注归纳偏差,使架构在不牺牲对任意输入表示支持的情况下,更好地对语音信号的时频结构进行建模。

英文摘要

Recently, time-frequency neural vocoders have been approaching the state-of-the-art quality of time-domain neural vocoders. Vocos is a notable example due to its efficiency, but its audio quality lags behind the time-domain vocoders and the reasons remain debated. Thus, in this study, we revisit Vocos from a phase reconstruction perspective. First, we quantify the gap between time-domain and time-frequency domain vocoders using bandlimited mel spectrograms as inputs. Later, via an ablation study, we verify the Vocos architecture is effective for magnitude modeling, but less so for phase. We then adapt the Vocos backbone to predict phase differences, a precursor for phase reconstruction, and identify 1D convolutional layers are hindering their accurate prediction. Our findings indicate that future research needs to focus on inductive biases that allow the architecture to better model the time-frequency structure of speech signals, without sacrificing the support for arbitrary input representations.

CommentsAccepted at IWAENC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑