arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RESTORE:基于音源分离的实时可操控音乐修复与带宽扩展

RESTORE: REal-time Steerable Music resTORation and bandwidth Extension via stem disentanglement

Meiying Chen, Benjamin R. Thompson, Michael C. Heilemann

arXiv 2609.28683首次发表:更新:

发表机构

University of Rochester(罗切斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RESTORE提出六源语义分解的音频修复框架,通过扩展HTDemucs实现实时交互式操控,在历史录音上降低FAD并实现高美学可操控性。

AI 中文摘要

神经音频修复方法通常被构建为从退化输入到单一干净输出的刚性映射,这强制决定了从音频中移除哪些内容,并可能向修复后的信号中添加不需要的内容。由于何为修复后的音频信号具有主观性,我们提出了RESTORE框架,该框架将音频修复表述为六源语义分解,以实现对修复过程的实时交互式用户控制。通过扩展预训练的HTDemucs骨干网络,单次前向传播即可将退化混合信号分离为人声、音乐、宽带嘶声、脉冲瞬态以及未建模残差,同时联合合成高频扩展。用户可通过调整音源增益来操控修复过程,确保生成内容保持隔离且可审计。与基线方法相比,RESTORE在多种历史录音上提升了音频质量,降低了Frechet音频距离(FAD)(VGGish为12.13;CLAP为0.92),并在单个GPU上以50倍实时速度实现了美学可操控性(Spearman rho大于0.91)。代码和音频样本可在该https URL获取。

英文摘要

Neural methods for audio restoration are typically framed as rigid mappings from degraded inputs to single clean outputs, enforcing decisions about what audio content is removed, and potentially adding unwanted content to the restored signal. Because what constitutes a restored audio signal is subjective, we introduce RESTORE, a framework that formulates audio restoration as a six-source semantic decomposition to allow for real-time interactive user control over the process. By expanding a pretrained HTDemucs backbone, a single forward pass disentangles a degraded mixture into vocals, music, broadband hiss, impulsive transients, and an unmodeled residual, while jointly synthesizing a high-frequency extension. Users may steer the restoration by adjusting stem gains, ensuring generative content remains isolated and auditable. RESTORE improves audio quality on diverse historical recordings compared to baselines,lowering Frechet Audio Distance (FAD) (12.13 VGGish; 0.92 CLAP) and delivering aesthetic steerability (Spearman rho greather than 0.91) at 50x real-time on a single GPU. Code and audio samples are available at https://melissachen15.github.io/restore-audio-demo.

CommentsCode and audio samples: https://melissachen15.github.io/restore-audio-demo

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑