arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12201cs.SD

低功耗音频DSP上的实时音乐源分离

Real-Time Music Source Separation on a Low-Power Audio DSP

Jianan Li, Li Liu, Ken Malsky, Gabby Yi

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对低功耗音频DSP,验证现有实时音乐源分离系统均不适用,提出在连续卷积上下文训练并加入门控复数FIR深度滤波器的新系统,在MUSDB18-HQ上达到4.70 dB cSDR,运行时间仅10.43 ms。

中文摘要 AI 辅助

实时音乐源分离已在桌面CPU和GPU上得到验证。是否有任何已发表的系统适配其目标嵌入式音频硬件?在商用音频DSP(2 MB SRAM,实测2.07 GMAC/s)上,没有系统能适配,且约束条件淘汰了不同模型:内存限制排除了16-51 M参数的TasNet/X-UMX系列,每帧计算量排除了RT-STT,其需要可用MAC速率的5.5倍。参数数量无法预测性能:权重复用范围从1倍到345倍。我们随后构建了一个适配的系统。在连续而非块填充的卷积上下文上训练被证明至关重要:一个在块-wise上得分为3.93 dB的模型,在逐帧处理时2秒内崩溃为静音。一个门控复数FIR深度滤波器增加了延迟旋钮,即使在严格因果的情况下也获得了0.38 dB的提升。该系统在MUSDB18-HQ上达到4.70 dB的cSDR,并在11.6 ms的跳数中运行10.43 ms,比不适配的系统落后0.5-0.7 dB。

英文摘要

Real-time music source separation is validated on desktop CPUs and GPUs. Does any published system fit the embedded audio hardware it targets? On a commercial audio DSP (2 MB SRAM, 2.07 GMAC/s measured), none does, and the constraints eliminate different models: memory rules out the 16-51 M parameter TasNet/X-UMX family, per-frame compute rules out RT-STT, needing 5.5x the available MAC rate. Parameter count predicts neither: weight reuse spans 1x to 345x. We then build one that fits. Training on continuous rather than block-padded convolution context proves essential: a model scoring 3.93 dB block-wise otherwise collapses to silence within 2 s frame-by-frame. A gated complex FIR deep filter adds a latency knob, gaining 0.38 dB even when strictly causal. It reaches 4.70 dB cSDR on MUSDB18-HQ and runs in 10.43 ms of an 11.6 ms hop, 0.5-0.7 dB behind systems that do not fit.

发表机构

  • Analog Devices, Inc.(亚德诺半导体公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑