arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23411cs.SDeess.AS

误识别还是抽象?重新思考声音事件识别的输出

Misrecognition or Abstraction? Rethinking Outputs of Sound Event Recognition

发表机构京都大学 · 东京大学
查看机构详情
  • Kyoto University(京都大学)
  • The University of Tokyo(东京大学)

机构由 AI 辅助整理,请以论文原文为准。

Naoya Tomida, Yuki Okamoto, Keisuke Imoto

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出结合声音类别、置信度和拟声词描述的输出表示,在不确定性下提升声音事件识别的信息性和沟通性,实验验证其性能与传统系统相当且更受偏好。

中文摘要 AI 辅助

传统的一般声音识别系统通常输出确定性的声音事件标签,隐含地假设可以从输入音频中正确识别目标声音类别。然而,在真实聆听情境中,声音事件类别并不总是清晰可辨的。人类听者仍然可能通过模糊的声音理解周围环境,而无需识别其确切的声音事件类别。这促使我们讨论在不确定性下应如何重新设计声音识别系统的输出。作为讨论的基础,本文提出了一种声音事件识别的输出表示,该表示结合了声音事件类别、其置信度分数以及声音的拟声词描述。所提出的表示保留了传统的基于类别的识别,同时提供了声学特征的额外拟声词描述,即使在类别预测不确定时,该描述也能保持信息量。使用ESC-50和ESC-50-Onomatopoeia的实验表明,所提出的方法实现了与传统仅识别系统相当的声音识别性能。此外,LLM-as-a-judge评估和主观聆听实验表明,所提出的输出优于基于声音事件标签的传统确定性输出,特别是在用于支持对周围环境的理解时。这些结果表明,这种输出表示可以使声音事件识别在不确定性下更具信息性和沟通性。

英文摘要

Conventional general sound recognition systems typically output deterministic sound event labels, implicitly assuming that the target sound class can be correctly identified from the input audio. However, in real listening situations, the sound event class is not always clearly identifiable. Human listeners may nevertheless understand their surroundings from an ambiguous sound without identifying its exact sound event class. This motivates a discussion of how the outputs of sound recognition systems should be redesigned under such uncertainty. As a basis for this discussion, this paper proposes an output representation for sound event recognition that combines a sound event class, its confidence score, and an onomatopoeic description of the sound. The proposed representation preserves conventional class-based recognition while providing an additional onomatopoeic description of acoustic characteristics that can remain informative even when the class prediction is uncertain. Experiments using ESC-50 and ESC-50-Onomatopoeia show that the proposed method achieves sound recognition performance comparable to that of a conventional recognition-only system. In addition, an LLM-as-a-judge evaluation and subjective listening experiments indicate that the proposed output is preferred over conventional deterministic outputs based on the sound event label, particularly when used to support understanding of the surrounding environment. These results suggest that such output representations can make sound event recognition more informative and communicative under uncertainty.

补充信息

↑