DiTAR+:面向鲁棒自回归扩散语音合成的双重优化
DiTAR+: Dual Optimization for Robust Autoregressive Diffusion Speech Synthesis
浏览论文内容
中文总结 AI 辅助
DiTAR+通过膨胀上下文采样与分层声学掩蔽双重优化,解决自回归扩散语音合成中的解码不稳定问题,降低词错误率并提升长语句说话人相似度。
中文摘要 AI 辅助
连续潜变量自回归扩散Transformer(AR-DiT)模型在零样本语音生成中展现出巨大潜力。然而,在合成长语句或复杂语言结构时,它们仍面临解码稳定性受限的问题。这种不稳定性主要源于扩散解码器内受限的历史感受野和声学惯性依赖,导致模型忽略语义条件。为解决这些挑战,我们提出DiTAR+,一个双重优化框架。首先,我们引入膨胀上下文采样(Dilated Context Sampling),在不违反物理时间连续性的前提下扩展宏观层面的历史感受野,从而防止累积误差传播。其次,我们提出分层声学掩蔽(Hierarchical Acoustic Masking),阻止浅层层关注声学前上下文,显式地将语义对齐与声学细节重建解耦。大量实验表明,我们的框架有效缓解了发音错误和语义幻觉,增强了在挑战性句子上的生成鲁棒性,并在整个长语句中保持了极高的说话人相似度。在语言上具有挑战性的ZH-Hard数据集上,DiTAR+将词错误率从12.478%降至9.893%;在25至35秒的扩展语句上,它将说话人相似度从0.741提升至0.759,同时将词错误率从2.778%降至2.173%,优于离散令牌和纯流匹配基线。
英文摘要
Continuous-latent Autoregressive Diffusion Transformer (AR-DiT) models have demonstrated immense potential in zero-shot speech generation. However, they still suffer from limited decoding stability when synthesizing long utterances or complex linguistic structures. This instability primarily stems from a restricted historical receptive field and an acoustic inertia dependency within the diffusion decoder, which causes the model to ignore semantic conditions. To address these challenges, we propose DiTAR+, a dual-optimization framework. First, we introduce Dilated Context Sampling to expand the macro-level historical receptive field without violating physical temporal continuity, thereby preventing cumulative error propagation. Second, we propose Hierarchical Acoustic Masking to prevent shallow layers from attending to acoustic pre-context, explicitly decoupling semantic alignment from acoustic detail reconstruction. Extensive experiments show that our framework effectively mitigates pronunciation errors and semantic hallucinations, enhances generation robustness on challenging sentences, and maintains exceptionally high speaker similarity throughout the entirety of long-form utterances. On the linguistically challenging ZH-Hard set, DiTAR+ reduces the word error rate from 12.478% to 9.893%, and on extended utterances of 25 to 35 seconds it improves speaker similarity from 0.741 to 0.759 while simultaneously lowering the word error rate from 2.778% to 2.173%, outperforming both discrete-token and pure flow-matching baselines.
发表机构
- Northwestern Polytechnical University(西北工业大学)
机构由 AI 辅助整理,请以论文原文为准。