arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将离线模型适配到流式上下文以进行音乐源分离

Adapting offline models to a streaming context for music source separation

Dylan Sechet, Marc Evrard, Matthieu Kowalski

arXiv 2609.35397首次发表:更新:

发表机构

Université Paris-Saclay; Inria; CNRS(巴黎萨克雷大学; 法国国家信息与自动化研究所; 法国国家科学研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文证明离线音乐源分离模型可通过滑动输入和输出实现低至23毫秒的流式运行,无需重新训练,且质量与专用实时模型相当,但计算效率仍待提升。

AI 中文摘要

实时音乐源分离必须满足两个约束:算法延迟的上限和计算成本的上限。离线分离器通常被排除在实时比较之外,或者被赋予等于其完整输入长度的延迟。我们表明,这个延迟是由输出被读取的位置决定的,而不是由分离器输入的长度决定的。因此,未经修改的离线模型可以在流式设置中运行,无需重新训练。在每一步中,输入滑动一个STFT跳数,并读出一个输出跳数。由此产生的延迟可以低至一个STFT跳数(23毫秒),并且计算成本不会随着延迟的缩小而增加。我们确定了一个理论上的模型相关延迟边界,低于该边界分离质量应急剧下降,并在三种架构上通过实验证实了这一点。在相同的算法延迟下,HT-Demucs和SCNet的流式现成检查点在分离质量方面与专门实时模型的已发表结果相匹配。流式模型的计算效率仍然远低:只有HT-Demucs在GPU上的运行速度快于实时。

英文摘要

Real-time music source separation must satisfy two constraints: a bound on algorithmic latency and a bound on computational cost. Offline separators are usually omitted from real-time comparisons or credited with a latency equal to their full input length. We show that this latency is set by where the output is read, not by the length of the separator's input. An unmodified offline model can therefore run in a streaming setting, without retraining. At each step, the input slides by one STFT hop, and one output hop is read out. The resulting latency can be as low as one STFT hop (23 ms), and the computational cost does not increase as latency shrinks. We identify a theoretical model-dependent latency boundary below which separation quality should drop steeply, and confirm this experimentally across three architectures. At equal algorithmic latency, streamed off-the-shelf checkpoints for HT-Demucs and SCNet match the published results of dedicated real-time models in terms of separation quality. Streamed models remain far less computationally efficient: only HT-Demucs runs faster than real time on a GPU.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑