Xiaomi-CocktailASR-1 技术报告
Xiaomi-CocktailASR-1 Technical Report
浏览论文内容
中文总结 AI 辅助
针对多说话人场景中鸡尾酒会问题,提出基于大语言模型的端到端目标说话人识别架构Xiaomi-CocktailASR-1,利用参考语音作为声纹提示直接转录目标语音,无需语音分离,具备弃权能力和思维链推理,在多个基准上达到最优性能。
中文摘要 AI 辅助
近年来,基于大语言模型(LLM)的自动语音识别(ASR)模型取得了显著进展,但它们普遍缺乏对多说话人场景的支持,其中鸡尾酒会问题仍是进一步推进ASR的关键瓶颈。现有的目标说话人ASR(TS-ASR)方法,包括带有说话人嵌入的端到端架构和最新的基于LLM的探索,都存在单说话人性能下降以及无法在目标说话人缺席时进行拒绝(即弃权(不执行))的问题。在本文中,我们提出了Xiaomi-CocktailASR-1,一种基于LLM的端到端TS-ASR架构。通过利用参考语音作为声纹提示,它直接转录目标说话人的语音,而无需语音分离。Xiaomi-CocktailASR-1在单说话人场景中保持了具有竞争力的性能,与主流ASR模型相当。它还具有负样本拒绝能力,当混合语音中不存在目标说话人时,输出空文本。此外,Xiaomi-CocktailASR-1支持思维链(CoT)推理模式,以提供明确的推理步骤。在各种合成和真实世界的多说话人基准上的大量实验表明,Xiaomi-CocktailASR-1实现了最先进的性能,通过统一架构有效解决了鸡尾酒会问题,平衡了多说话人和单说话人识别准确性以及拒绝能力。
英文摘要
Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party problem remains a critical bottleneck for further advancing ASR. Existing TS-ASR methods, including end-to-end architectures with speaker embeddings and latest LLM-based explorations suffer from degraded single-speaker performance and the inability to reject when the target speaker is absent. In this paper, we propose Xiaomi-CocktailASR-1, an LLM-based end-to-end TS-ASR architecture. By utilizing reference speech as voiceprint prompts, it directly transcribes the target speaker's speech without requiring speech separation. Xiaomi-CocktailASR-1 maintains competitive performance in single-speaker scenarios, comparable to mainstream ASR models. It also features a negative sample rejection capability, outputting empty text when the target speaker is absent from the mixed speech. Additionally, Xiaomi-CocktailASR-1 supports a Chain-of-Thought (CoT) reasoning mode to provide explicit reasoning steps. Extensive experiments on various synthetic and real-world multispeaker benchmarks demonstrate that Xiaomi-CocktailASR-1 achieves state-of-the-art performance, effectively addressing the cocktail party problem through a unified architecture that balances multispeaker and single-speaker recognition accuracy, along with rejection capability.
发表机构
- Xiaomi Inc.(小米公司)
机构由 AI 辅助整理,请以论文原文为准。