发表机构
Inria; AMIAD(法国国家信息与自动化研究所; AMIAD)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出CARES基准,包含10,000个受控双说话者场景,定义声音显著性为说话者反应,并评测六个音频语言模型,发现它们能识别声音但忽略说话者反应。
AI 中文摘要
自动音频场景描述将录音转化为对情境的文字叙述。一个难点在于决定应保留音频中的哪些元素,因为描述无法涵盖所有内容。标注者对此意见不一,使得获取真实标注变得困难。在本工作中,我们首先定义真实标注,然后生成数据。我们聚焦于音频事件,并用一个简单规则定义声音显著性:当说话者对其有可听见的反应时,该声音即为显著。为了规模和多样性,一组受控场景固定了真实标注,并由语言模型编写对话。由此产生的语料库CARES包含10,000个双说话者场景。随后,我们在三个任务上对六个音频语言模型进行基准测试:识别场景、标记存在的声音以及分类反应。我们表明,这些模型能听到声音,但错过了说话者对其的反应。
英文摘要
Automatic audio scene description turns a recording into a text account of a situation. One difficulty is deciding which elements of the audio should be kept, since a description cannot include them all. Annotators disagree about this, making a ground truth hard to obtain. In this work, we first define the ground truth, then generate the data. We focus on audio events and define sound salience with a simple rule: a sound is salient when a speaker audibly reacts to it. For scale and variety, a controlled set of scenarios fixes the ground truth, and a language model writes the dialogues. The resulting corpus, CARES, contains 10,000 two-speaker scenes. We then benchmark six audio-language models on three tasks: identifying the scene, tagging the sounds present, and classifying reactions. We show that these models hear the sounds but miss how the speakers react to them.
CommentsSubmitted to ICASSP 2027