VibeVoice-ASR-BitNet技术报告
VibeVoice-ASR-BitNet Technical Report
- Microsoft Research(微软研究院)
- Shanghai Jiao Tong University(上海交通大学)
- Fudan University(复旦大学)
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
介绍VibeVoice-ASR-BitNet,针对边缘CPU实时推理优化。采用异构量化及渐进式量化感知训练,在ggml框架实现自定义内核和运算符,可比模型大小下速度快1.6 - 2.3倍,准确性略降。
AI中文摘要:
我们展示了VibeVoice-ASR-BitNet,它是VibeVoice-ASR的压缩变体,专为在边缘CPU上进行实时推理而优化。我们针对每个阶段的计算特性应用异构量化:VAE声学分词器使用带内核融合和SIMD优化的全流水线INT8量化(I8_S),而自回归语言模型采用BitNet风格的三元权重(I2_S)。为在激进压缩下保持准确性,我们采用渐进式量化感知训练策略。在推理时,我们在ggml框架内针对ARM和x86平台实现自定义SIMD内核和融合运算符,使用仅3个CPU线程就能实现实时识别,RTF<1。在可比模型大小(约1.6GB)下,VibeVoice-ASR-BitNet比该http URL快1.6-2.3倍,与FP16基线相比,准确性仅略有下降。
英文摘要:
We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computational characteristics of each stage: the VAE acoustic tokenizer uses full-pipeline INT8 quantization (I8_S) with kernel fusion and SIMD optimization, while the autoregressive language model adopts BitNet-style ternary weights (I2_S). To preserve accuracy under aggressive compression, we employ a progressive quantization-aware training strategy. For inference, we implement custom SIMD kernels and fused operators within the ggml framework targeting both ARM and x86 platforms, achieving real-time recognition (RTF < 1) on low-thread-count CPUs. VibeVoice-ASR-BitNet is 1.6--2.3x faster than Whisper.cpp at comparable model sizes (~1.6 GB), with only modest accuracy degradation compared to the FP16 baseline.