arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11274cs.SDcs.CLeess.AS

Xiaomi-CocktailASR-1 技术报告

Xiaomi-CocktailASR-1 Technical Report

Yiru Zhang, Hang Su, Lichun Fan, Ying Zeng, Chang Liu, Yifeng Wang, Yuquan Liang, Tao Li, Lian Li, Wenhao Yang, Jian Luan, Cong Zou, Heng Qu

首次发表
浏览论文内容

中文总结 AI 辅助

针对多说话人场景中鸡尾酒会问题,提出基于大语言模型的端到端目标说话人识别架构Xiaomi-CocktailASR-1,利用参考语音作为声纹提示直接转录目标语音,无需语音分离,具备弃权能力和思维链推理,在多个基准上达到最优性能。

中文摘要 AI 辅助

近年来,基于大语言模型(LLM)的自动语音识别(ASR)模型取得了显著进展,但它们普遍缺乏对多说话人场景的支持,其中鸡尾酒会问题仍是进一步推进ASR的关键瓶颈。现有的目标说话人ASR(TS-ASR)方法,包括带有说话人嵌入的端到端架构和最新的基于LLM的探索,都存在单说话人性能下降以及无法在目标说话人缺席时进行拒绝(即弃权(不执行))的问题。在本文中,我们提出了Xiaomi-CocktailASR-1,一种基于LLM的端到端TS-ASR架构。通过利用参考语音作为声纹提示,它直接转录目标说话人的语音,而无需语音分离。Xiaomi-CocktailASR-1在单说话人场景中保持了具有竞争力的性能,与主流ASR模型相当。它还具有负样本拒绝能力,当混合语音中不存在目标说话人时,输出空文本。此外,Xiaomi-CocktailASR-1支持思维链(CoT)推理模式,以提供明确的推理步骤。在各种合成和真实世界的多说话人基准上的大量实验表明,Xiaomi-CocktailASR-1实现了最先进的性能,通过统一架构有效解决了鸡尾酒会问题,平衡了多说话人和单说话人识别准确性以及拒绝能力。

英文摘要

Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party problem remains a critical bottleneck for further advancing ASR. Existing TS-ASR methods, including end-to-end architectures with speaker embeddings and latest LLM-based explorations suffer from degraded single-speaker performance and the inability to reject when the target speaker is absent. In this paper, we propose Xiaomi-CocktailASR-1, an LLM-based end-to-end TS-ASR architecture. By utilizing reference speech as voiceprint prompts, it directly transcribes the target speaker's speech without requiring speech separation. Xiaomi-CocktailASR-1 maintains competitive performance in single-speaker scenarios, comparable to mainstream ASR models. It also features a negative sample rejection capability, outputting empty text when the target speaker is absent from the mixed speech. Additionally, Xiaomi-CocktailASR-1 supports a Chain-of-Thought (CoT) reasoning mode to provide explicit reasoning steps. Extensive experiments on various synthetic and real-world multispeaker benchmarks demonstrate that Xiaomi-CocktailASR-1 achieves state-of-the-art performance, effectively addressing the cocktail party problem through a unified architecture that balances multispeaker and single-speaker recognition accuracy, along with rejection capability.

发表机构

  • Xiaomi Inc.(小米公司)

机构由 AI 辅助整理,请以论文原文为准。

相关深度报道

↑