发表机构
University of Chinese Academy of Sciences; Microsoft Research; Shanghai Jiao Tong University(中国科学院大学; 微软研究院; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有统一带说话人属性ASR模型难以满足实时语音助手低延迟需求的问题,提出基于LLM的VibeVoice-ASR-Streaming,其在转录准确率和说话人属性评估中表现优异,并发布了对应模型权重与推理代码。
AI 中文摘要
传统带说话人属性的自动语音识别(ASR)系统将ASR与说话人 diarization(说话人 diarization 指说话人分割与识别任务)视为两个独立任务。近期,VibeVoice-ASR 等端到端模型将这两个任务统一到单个模型中,但现有统一模型仍主要支持离线识别,难以满足实时语音助手和智能体的低延迟需求。为解决该问题,我们提出 VibeVoice-ASR-Streaming,它是首批基于大语言模型(LLM)的端到端流式带说话人属性的 ASR 方法之一。该模型交错处理固定大小的音频块、少量前瞻音频及前序文本,可在语音输入时同步输出“谁在说什么”,无需单独的 diarization 阶段。转录准确率方面,我们的7B模型在5个评估数据集上取得最低的平均 WER(词错误率)/CER(字符错误率);说话人属性方面,它在13个评估设置中的12个取得最佳成绩或并列最佳成绩。我们发布了1.5B和7B的模型权重及推理代码。
英文摘要
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, a small amount of lookahead audio and previous text. This allows the model to produce ''who said what'' as speech arrives, without a separate diarization stage. For transcription accuracy, our 7B model achieves the lowest average WER/CER across five evaluation sets. For speaker attribution, it achieves the best or tied-best on 12 of 13 evaluation settings. We release the 1.5B and 7B model weights together with inference code.