arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21075cs.SDcs.CLeess.AS

VibeVoice-ASR-BitNet技术报告

VibeVoice-ASR-BitNet Technical Report

  • Microsoft Research(微软研究院)
  • Shanghai Jiao Tong University(上海交通大学)
  • Fudan University(复旦大学)
  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

Songchen Xu, Ting Song, Shaohan Huang, Zhiliang Peng, Yan Xia, Yujie Tu, Xin Huang, Xun Wu, Wenhui Wang, Yaoyao Chang, Jianwei Yu, Li Dong, Furu Wei

AI总结:

介绍VibeVoice-ASR-BitNet,针对边缘CPU实时推理优化。采用异构量化及渐进式量化感知训练,在ggml框架实现自定义内核和运算符,可比模型大小下速度快1.6 - 2.3倍,准确性略降。

AI中文摘要:

我们展示了VibeVoice-ASR-BitNet,它是VibeVoice-ASR的压缩变体,专为在边缘CPU上进行实时推理而优化。我们针对每个阶段的计算特性应用异构量化:VAE声学分词器使用带内核融合和SIMD优化的全流水线INT8量化(I8_S),而自回归语言模型采用BitNet风格的三元权重(I2_S)。为在激进压缩下保持准确性,我们采用渐进式量化感知训练策略。在推理时,我们在ggml框架内针对ARM和x86平台实现自定义SIMD内核和融合运算符,使用仅3个CPU线程就能实现实时识别,RTF<1。在可比模型大小(约1.6GB)下,VibeVoice-ASR-BitNet比该http URL快1.6-2.3倍,与FP16基线相比,准确性仅略有下降。

英文摘要:

We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computational characteristics of each stage: the VAE acoustic tokenizer uses full-pipeline INT8 quantization (I8_S) with kernel fusion and SIMD optimization, while the autoregressive language model adopts BitNet-style ternary weights (I2_S). To preserve accuracy under aggressive compression, we employ a progressive quantization-aware training strategy. For inference, we implement custom SIMD kernels and fused operators within the ggml framework targeting both ARM and x86 platforms, achieving real-time recognition (RTF < 1) on low-thread-count CPUs. VibeVoice-ASR-BitNet is 1.6--2.3x faster than Whisper.cpp at comparable model sizes (~1.6 GB), with only modest accuracy degradation compared to the FP16 baseline.

补充信息

↑