发表机构
JKU(林茨约翰内斯开普勒大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究严格因果和低延迟约束下的实时音乐增强,采用紧凑因果网络并与多种模型比较,结果显示改进因多种因素而异,主要贡献是给出基准和分析,强调实时音乐增强虽可行,但稳健改进需多方面考量。
AI 中文摘要
音乐录音和直播流常受噪声、混响、频谱不平衡或伪影影响,导致收听质量下降。语音增强已发展成熟,而音乐增强因音乐信号的复杂性尚不完善。本文研究严格因果和低延迟约束下的实时音乐增强。围绕从声学和生产导向的退化中恢复预期混音制定任务,采用紧凑因果网络并与多种模型比较。结果表明所有因果模型运行速度快于实时,但改进因数据集、退化类型和指标而异,不加区分的增强可能恶化输入。主要贡献是一个基准和分析,即实时音乐增强可行,但稳健改进需考虑退化感知建模等多方面。
英文摘要
Music recordings and live streams are often affected by noise, reverberation, spectral imbalances, or artifacts that degrade listening quality. While speech enhancement has matured into a well-defined research area, music enhancement is less established because musical signals combine overlapping sources, wide bandwidths, strong dynamics, and intentional production effects. We study real-time music enhancement under strict causal and low-latency constraints. We formulate the task around recovery of the intended produced mix from acoustic and production-oriented degradations, adapt compact causal networks to music, and compare speech-derived real-time baselines, an external music-denoising model, an offline restoration reference, and a music-specific MusicFilterNet-MS variant. On the tested hardware, all causal models run faster than real time, but improvements depend strongly on the dataset, degradation type, and metric family; under several objective criteria, indiscriminate enhancement can worsen the degraded input. The main contribution is therefore a benchmark and an analysis rather than a universal best model: real-time music enhancement is feasible, but robust improvement requires degradation-aware modeling, stereo-aware processing, identity-preserving correction, and evaluation beyond a single objective score.