音乐源分离:通过音轨发现
Music Source Separation via Stem Discovery
浏览论文内容
中文总结 AI 辅助
本文提出MuS3D,一种基于音频查询的迭代式音乐源分离框架,通过自动发现混合中的活动源,匹配手动查询基线并超越文本模型,证明音频查询是有效且可自动化的分离接口。
中文摘要 AI 辅助
音乐源分离(MSS)方法旨在从音乐混合中提取音轨,这在例如卡拉OK、音乐混音和教学应用中非常重要。虽然早期关于MSS系统的研究主要由针对狭窄的通用音轨集合的模型主导,但最近已有尝试支持更广泛的源定义。其中一种方法使用音频查询,基于声音本身对期望的分离目标提供直接且描述性的控制。然而,基于查询的分离仍然很繁琐,因为需要提供与混合中包含的源特征匹配的音频示例。本文提出了音乐源分离:通过音轨发现(MuS3D),一种基于查询的源分离框架,它从混合中迭代地发现活动源。在正确检测到的源上,我们的模型匹配手动查询的基线,并超越最先进的基于文本的模型。主观评估表明,编码伪影是当前生成式分离中的限制因素。研究结果表明,基于音频的查询表示为源分离提供了一种有效且可自动化的接口。
英文摘要
Music source separation (MSS) methods aim to extract stems from music mixtures, which is important, for example, in karaoke, music remixing, and pedagogical applications. While earlier research on MSS systems has been dominated by models targeting narrow sets of general stems, there have recently been attempts to support broader source definitions. One of these methods uses audio queries to provide direct and descriptive control over the desired separation targets based on the sound itself. However, query-based separation remains cumbersome due to the need to provide audio examples with features matching the sources contained within mixtures. This paper proposes Music Source Separation via Stem Discovery (MuS3D), a query-based source separation framework that iteratively discovers active sources from the mixture. On correctly detected sources, our model matches manually queried baselines and surpasses state-of-the-art text-based models. Subjective evaluation indicates encoding artifacts as the limiting factor in current generative separation. The findings suggest that audio-based query representations offer an effective and automatable interface for source separation.
发表机构
- Aalto University(阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。