arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12703eess.AS

音频分割:一种探索未知事件类别的音频记录的新范式

Audio Diarization: A New Paradigm for Exploring Audio Recordings with Unknown Event Classes

Alexander Werning, Reinhold Haeb-Umbach

首次发表
浏览论文内容

中文总结 AI 辅助

针对未知环境中声音事件类别未知的情况,提出音频分割任务,先定位声音事件,再分类。介绍了如何调整说话者分割系统用于音频分割及评估设置,该系统能检测新声音,与封闭集检测系统性能相似。

中文摘要 AI 辅助

我们提出了一项新任务——音频分割。其动机在于,在诸如未知环境中的音频监测等应用中,最初要识别的声音事件类别是未知的。对于这种情况,我们建议首先及时定位相关声音事件,然后在第二步中对其进行分类,例如通过与已知事件类别进行比较。本论文致力于第一步,即我们所称的音频分割,因为它类似于多说话者对话语音处理中先于并简化第二步语音识别的说话者分割阶段。在本论文中,我们将音频分割定义为在无用户提示的情况下,为一组开放类检测具有重叠的声音事件的起始和结束时间。我们展示了如何调整说话者分割系统以用于音频分割,并提出了一种评估设置。与封闭集声音事件检测系统相比,所提出的系统在具有检测新声音的额外能力的情况下实现了相似的性能。

英文摘要

We propose a new task, audio diarization. The motivation is that there are applications, such as audio monitoring in an unknown environment, where initially the sound event classes to be recognized are unknown. For such a scenario, we propose to first localize in time relevant sound events and to classify them, e.g., by comparing with known event classes, in a second step. This contribution is dedicated to the first step, which we call audio diarization, as it is reminiscent of the speaker diarization stage that precedes and simplifies the second stage, speech recognition, in multi-talker conversational speech processing. In this contribution, we define audio diarization as detecting onset and offset times of sound events with overlap for an open set of classes and without user prompts. We show how a speaker diarization system can be adjusted for audio diarization and propose an evaluation setup. Compared to a closed-set sound event detection system, the proposed system achieves similar performance with the additional ability to detect novel sounds.

补充信息

↑