从语音到交互:鸡尾酒会场景下多模态系统的分析
From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios
浏览论文内容
中文总结 AI 辅助
本研究分析了CHiME-9 MCoRec任务中不同多模态系统解决鸡尾酒会场景的策略,发现最优系统实现57%相对错误减少,且高语音重叠并非性能差异的主要原因。
中文摘要 AI 辅助
人类具备在自发非正式对话中,选择性关注特定说话者、同时过滤周边对话干扰语音的卓越能力。这种“鸡尾酒会”场景对语音识别系统仍构成严峻挑战。CHiME-9 MCoRec任务提供了一个测试平台,要求系统从视听输入中识别说话者群组并转录每个说话者的对话。本研究分析了一系列代表不同鸡尾酒会场景解决方向的多样系统,其中最优系统实现了高达57%的相对错误减少。我们确定了三种主要策略:(1)显式或隐式的视听目标语音分离;(2)针对每个目标说话者的改进视听语音识别;(3)利用大型语言模型将说话者分组为对话并增强对话一致性。我们的分析表明,这些方向解决了鸡尾酒会问题的互补失败模式,且仅高语音重叠无法解释性能差异,挑战了“重叠是鸡尾酒会识别主要困难来源”的普遍假设。
英文摘要
Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby conversations. This "cocktail party" scenario still presents severe challenges to speech recognition systems. The CHiME-9 MCoRec task provides a testbed where systems must recognize groups of speakers and transcribe each of their conversations from audio-visual input. In this work, we analyze a diverse set of systems, representing different design directions for addressing the cocktail-party scenario, where the best system achieves up to 57% relative error reduction. We identify three main strategies: (1) explicit or implicit audio-visual target speech separation, (2) improved audio-visual speech recognition for each target speaker, and (3) the use of large language models to group speakers into conversations and enhance conversational consistency. Our analysis shows that these directions address complementary failure modes of the cocktail-party problem, and that high speech overlap alone does not explain performance differences, challenging the common assumption that overlap is the primary source of difficulty in cocktail-party recognition.
发表机构
- Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。