arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LiteCASS:用于实时立体声电影音频源分离的轻量级端到端网络

LiteCASS: A Lightweight End-to-End Network for Real-Time Stereo Cinematic Audio Source Separation

Yuanxin Guo, Qiang Ji, Mengmei Liu, Yuhan Lv, Ningning Pan, Gongping Huang

arXiv 2609.23453首次发表:更新:

发表机构

Southwestern University of Finance and Economics; Xiaomi Automobile Co., Ltd.; Wuhan University(西南财经大学; 小米汽车有限公司; 武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有电影音频源分离方法参数多、依赖GPU且仅支持单声道的问题,提出轻量级端到端网络LiteCASS,结合STFT子带重排与双U-Net实现实时立体声分离,以1.06M参数和0.72G MACs在DnR v3立体声扩展上取得最高平均SI-SDR。

AI 中文摘要

电影音频源分离(CASS)将配乐分解为对白、音乐和音效(SFX)主干。然而,现有的CASS方法存在两个关键局限:它们依赖参数规模庞大的网络架构和GPU级硬件,限制了其在实时和资源受限场景中的使用;并且它们绝大多数是为单声道信号设计的,立体声场景在很大程度上尚未被探索。我们提出了LiteCASS,据我们所知,这是首个用于实时立体声CASS的轻量级端到端网络。LiteCASS将确定性的STFT子带重排与两个联合训练的紧凑型U-Net相结合:第一个提取对白,第二个从预测的非语音成分中分离音乐和音效。多任务波形域L1损失监督所有主干。在DnR v3的空间化立体声扩展上,LiteCASS-K8仅使用1.06M参数和每秒0.72G MACs,同时在所比较的CASS基线中实现了最高的平均SI-SDR。

英文摘要

Cinematic audio source separation (CASS) decomposes a soundtrack into dialogue, music, and sound-effects (SFX) stems. Existing CASS methods, however, suffer from two critical limitations: they rely on heavily parameterized network architectures and GPU-class hardware, limiting their use in real-time and resource-constrained scenarios, and they are overwhelmingly designed for monaural signals, leaving the stereo scenario largely unexplored. We present LiteCASS, to our knowledge the first lightweight end-to-end network for real-time stereo CASS. LiteCASS combines deterministic STFT subband rearrangement with two jointly trained compact U-Nets: the first extracts dialogue, and the second separates music and SFX from the predicted non-speech component. A multi-task waveform-domain L1 loss supervises all stems. On a spatialized stereo extension of DnR v3, LiteCASS-K8 uses only 1.06M parameters and 0.72G MACs per second, while achieving the highest averaged SI-SDR among the compared CASS baselines.

Comments5 pages, 2 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑