arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于双时频谱表示的神经音乐增强方法(用于预测与判别)

Neural Music Enhancement with Dual Time-Frequency Spectral Representations for Prediction and Discrimination

Fei Liu, Yang Ai, Zhen-Hua Ling

arXiv 2609.03357首次发表:更新:

发表机构

National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China(中国科学技术大学国家语音及语言信息处理工程研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对非专业音乐录音的噪声混响问题,提出基于双时频谱的DSME模型,采用STFT生成与CQT判别,引入色度谱损失,实验验证其增强效果优于基线。

AI 中文摘要

网上分享的非专业音乐录音常存在背景噪声和混响,降低了感知质量并限制了其复用。本文提出DSME,一种基于双时频谱表示的音乐增强模型。在生成对抗框架内,DSME使用短时傅里叶变换(STFT)谱进行生成,使用恒Q变换(CQT)谱进行判别。利用STFT的固定窗口、可逆性和可预测性,生成器从降质输入估计纯净的幅度-相位谱,并通过逆STFT重构波形。利用CQT的对数频率、与音乐八度对齐的可变窗口结构,我们设计了八度分段的CQT判别器。我们还引入了色度谱损失以强调音高与谐波一致性。实验表明,DSME在客观与主观测试中均优于基线模型,验证了双谱方法的有效性。

英文摘要

Non-professional music recordings shared online often suffer from background noise and reverberation, degrading perceived quality and limiting reuse. This paper proposes DSME, a music enhancement model based on dual time-frequency spectral representations. Within a generative adversarial framework, DSME uses short-time Fourier transform (STFT) spectra for generation and constant-Q transform (CQT) spectra for discrimination. Leveraging STFT's fixed window, invertibility, and predictability, the generator estimates clean amplitude-phase spectra from degraded inputs and reconstructs waveforms via inverse STFT. Exploiting CQT's log-frequency, variable-window structure aligned with musical octaves, we design an octave-segmented CQT discriminator. We also introduce a chroma-spectrum loss to emphasize pitch and harmonic consistency. Experiments show DSME outperforms baselines in objective and subjective tests, validating the effectiveness of the dual-spectrum approach.

CommentsAccepted by ISCSLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑