arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VibeVoice-ASR-Streaming 技术报告

VibeVoice-ASR-Streaming Technical Report

Yujie Tu, Zhiliang Peng, Jianwei Yu, Li Dong, Songchen Xu, Yaoyao Chang, Wenhui Wang, Zilong Wang, Zehua Wang, Yan Xia, Ruibin Yuan, Jiajun Zhang, Xie Chen, Furu Wei

arXiv 2609.02812首次发表:更新:

发表机构

University of Chinese Academy of Sciences; Microsoft Research; Shanghai Jiao Tong University(中国科学院大学; 微软研究院; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有统一带说话人属性ASR模型难以满足实时语音助手低延迟需求的问题,提出基于LLM的VibeVoice-ASR-Streaming,其在转录准确率和说话人属性评估中表现优异,并发布了对应模型权重与推理代码。

AI 中文摘要

传统带说话人属性的自动语音识别(ASR)系统将ASR与说话人 diarization(说话人 diarization 指说话人分割与识别任务)视为两个独立任务。近期,VibeVoice-ASR 等端到端模型将这两个任务统一到单个模型中,但现有统一模型仍主要支持离线识别,难以满足实时语音助手和智能体的低延迟需求。为解决该问题,我们提出 VibeVoice-ASR-Streaming,它是首批基于大语言模型(LLM)的端到端流式带说话人属性的 ASR 方法之一。该模型交错处理固定大小的音频块、少量前瞻音频及前序文本,可在语音输入时同步输出“谁在说什么”,无需单独的 diarization 阶段。转录准确率方面,我们的7B模型在5个评估数据集上取得最低的平均 WER(词错误率)/CER(字符错误率);说话人属性方面,它在13个评估设置中的12个取得最佳成绩或并列最佳成绩。我们发布了1.5B和7B的模型权重及推理代码。

英文摘要

Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, a small amount of lookahead audio and previous text. This allows the model to produce ''who said what'' as speech arrives, without a separate diarization stage. For transcription accuracy, our 7B model achieves the lowest average WER/CER across five evaluation sets. For speaker attribution, it achieves the best or tied-best on 12 of 13 evaluation settings. We release the 1.5B and 7B model weights together with inference code.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑