arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过解耦CLAP查询优化与自动化数据引擎改进多类音频源分离的研究

A Study on Improving Multi-class Audio Source Separation Via Decoupled CLAP Query Optimization and an Automated Data Engine

Amirhossein Hajavi, Hanhee Lee, Pushya Jain, Sky Qiao, Emmanuel Ko, Yuanhao Yu, Irina Kezele

arXiv 2610.10025首次发表:更新:

AI 中文总结

针对语言查询音频源分离适配难题,提出自动化数据引擎与两阶段CLAP控制信号优化框架,在七类声音上显著提升分离性能,主观评测优于基线与商业模型。

AI 中文摘要

语言查询音频源分离(LASS)能够利用自然语言提取任意声源。然而,由于训练数据噪声大以及基于CLAP的控制信号语义覆盖有限,将LASS模型适配到特定应用场景的声音类别颇具挑战。我们提出一个框架,包含用于训练数据整理的自动化数据引擎和用于类别特定CLAP控制信号的两阶段优化过程。我们在七个声音类别上的客观评估表明,数据精炼和控制信号优化持续提升了源分离性能。17名参与者参与的主观评估进一步证明,使用优化控制信号训练的模型在感知质量上优于基线模型及类似的商业模型。

英文摘要

Language-queried audio source separation (LASS) enables extracting any sound source using natural language. However, adapting LASS models to application-specific sound classes is challenging due to noisy training data and limited semantic coverage of the CLAP-based control signals. We propose a framework comprised of an automated data engine for training-data curation and a two-stage optimization process for class-specific CLAP control signals. Our objective evaluations across seven sound classes show that data refinement and control signal optimization consistently improve source separation performance. Subjective evaluation with 17 participants further demonstrates perceptual improvements of the model trained with optimized control signals over baseline and similar commercial models.

CommentsProject page: https://amhajavi.github.io/AudioSourceSeparation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑