arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于本体的目标声音提取

Ontology-based Target Sound Extraction

Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai, Naohiro Tawara, Shoko Araki

arXiv 2609.00752首次发表:更新:

AI 中文总结

该研究提出基于本体的目标声音提取新任务,构建AudioSet衍生本体的可学习类别嵌入表,用CPCC损失正则化,实验验证了TSE系统训练中考虑本体结构的益处。

AI 中文摘要

目标声音提取(TSE)旨在根据语义查询从混合声音中分离出感兴趣的声源。现有TSE系统依赖于与单个声音类别绑定的固定类别表示,限制了其处理自然组织环境声音的层级关系的能力。本文提出了基于本体的TSE,这是一种新的任务形式,单个模型可提取在声音本体任意层级查询的声音,从猫、狗等细粒度叶类到动物等高级类别。我们提出了在AudioSet衍生本体的所有节点上定义的可学习类别嵌入表,并用共表型相关系数(CPCC)损失进行正则化,该损失使嵌入距离与本体树中的最短路径距离对齐。我们对不同方法的实验表明,训练TSE系统时考虑本体结构具有益处。

英文摘要

Target sound extraction (TSE) aims to isolate a sound source of interest from a mixture, given a semantic query. Existing TSE systems are conditioned on fixed class representations tied to individual sound categories, limiting their ability to handle the hierarchical relationships that naturally organize environmental sounds. In this paper, we introduce ontology-based TSE, a new task formulation in which a single model extracts sounds queried at any level of a sound ontology, from fine-grained leaf classes such as cat and dog to high-level categories such as animal. We propose a learnable class embedding table defined over all nodes of an AudioSet-derived ontology, regularized with a Cophenetic Correlation Coefficient (CPCC) loss that aligns embedding distances with shortest-path distances in the ontology tree. Our experiments across different approaches show the benefit of considering the ontology structure when training TSE systems.

CommentsAccepted for publication at IWAENC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑