发表机构
Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出基于解耦时间深度扩散变换器的内存高效音频合成架构,通过特定设计将语义音频令牌转换为RVQ表示,用单个深度解码器和因果滑动窗口注意力降低内存复杂度,在AMX上实现高效合成,提升音频质量。
AI 中文摘要
Siri Expressive Voices利用苹果最强大的设备基础模型AFM 3 Core Advanced实时在设备上合成丰富、可配置的语音。本文介绍了实现该功能的内存高效音频合成架构:一个去令牌化器,它在苹果矩阵协处理器(AMX)严格的计算和内存预算内,将基础模型发出的语义音频令牌转换为高保真音频。我们通过三分量设计、流编码器、时间解码器和深度解码器将语义音频令牌转换为残差向量量化(RVQ)表示,系统地解耦时间和深度处理。具有Diffusion Transformer(DiT)风格阶段条件的单个可重用深度解码器自回归生成所有RVQ级别,取代了先前多解码器架构中的专用每层解码器,而具有固定窗口键值缓存的因果滑动窗口注意力产生与序列长度无关的恒定内存复杂度。在AMX上部署时,去令牌化器每生成一步大约持续10毫秒,比实时快约16倍,峰值运行时内存仅为21MB,设备资产为329MB,能够连续流式合成20 - 320秒的音频。这种恒定的小占用空间取代了传统基于变压器和GAN方法的线性和二次内存扩展。消融研究验证了关键架构组件,音频质量评估证实该架构在保持合成保真度的同时比现有方法提高了效率。在AFM 3 Core Advanced内以10亿参数激活大小运行时,与先前的设备文本转语音系统相比,整体平均意见得分提高了+0.28(4.15对3.87),对话语音提高了+0.42(4.24对3.82)。
英文摘要
Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple's most powerful on-device foundation model. This work presents the memory-efficient audio synthesis architecture behind that capability: a detokenizer that converts the semantic audio tokens emitted by the foundation model into high-fidelity audio within the tight compute and memory budget of the Apple Matrix Coprocessor (AMX). We convert semantic audio tokens to a residual vector quantization (RVQ) representation with a three-component design, a streaming encoder, a temporal decoder, and a depth decoder, that systematically decouples temporal and depth processing. A single reusable depth decoder with Diffusion Transformer (DiT)-style stage conditioning generates all RVQ levels autoregressively, replacing the dedicated per-level decoders of prior multi-decoder architectures, while causal sliding window attention with fixed-window key-value caching yields constant memory complexity independent of sequence length. Deployed on the AMX, the detokenizer sustains roughly 10 ms per generation step, about 16x faster than real time, with a peak runtime memory of only 21 MB and 329 MB of on-device assets, enabling continuous streaming synthesis of 20-320 seconds of audio. This constant, small footprint replaces the linear and quadratic memory scaling of conventional transformer- and GAN-based approaches. Ablation studies validate the key architectural components, and audio quality assessment confirms that the architecture maintains synthesis fidelity while achieving efficiency gains over existing methods. Operating at a 1-billion-parameter activation size within AFM 3 Core Advanced, it improves Mean Opinion Score by +0.28 overall (4.15 vs. 3.87) and by +0.42 on conversational speech (4.24 vs. 3.82) over the prior on-device text-to-speech system.
Comments11 pages, ICASSP