arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CARES:一个受控的说话者对声音反应的合成基准

CARES: A Controlled Synthetic Benchmark of Speaker Reactions to Sound

Marcel Gibier, Thomas Thebaud, Olivier Boëffard, Jean-François Bonastre

arXiv 2610.10208首次发表:更新:

发表机构

Inria; AMIAD(法国国家信息与自动化研究所; AMIAD)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出CARES基准,包含10,000个受控双说话者场景,定义声音显著性为说话者反应,并评测六个音频语言模型,发现它们能识别声音但忽略说话者反应。

AI 中文摘要

自动音频场景描述将录音转化为对情境的文字叙述。一个难点在于决定应保留音频中的哪些元素,因为描述无法涵盖所有内容。标注者对此意见不一,使得获取真实标注变得困难。在本工作中,我们首先定义真实标注,然后生成数据。我们聚焦于音频事件,并用一个简单规则定义声音显著性:当说话者对其有可听见的反应时,该声音即为显著。为了规模和多样性,一组受控场景固定了真实标注,并由语言模型编写对话。由此产生的语料库CARES包含10,000个双说话者场景。随后,我们在三个任务上对六个音频语言模型进行基准测试:识别场景、标记存在的声音以及分类反应。我们表明,这些模型能听到声音,但错过了说话者对其的反应。

英文摘要

Automatic audio scene description turns a recording into a text account of a situation. One difficulty is deciding which elements of the audio should be kept, since a description cannot include them all. Annotators disagree about this, making a ground truth hard to obtain. In this work, we first define the ground truth, then generate the data. We focus on audio events and define sound salience with a simple rule: a sound is salient when a speaker audibly reacts to it. For scale and variety, a controlled set of scenarios fixes the ground truth, and a language model writes the dialogues. The resulting corpus, CARES, contains 10,000 two-speaker scenes. We then benchmark six audio-language models on three tasks: identifying the scene, tagging the sounds present, and classifying reactions. We show that these models hear the sounds but miss how the speakers react to them.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑