AI 中文总结
AlignDPO通过将CTC对齐项融入DPO,实现中等对齐锐度,显著降低仅解码器TTS的内容幻觉率至0.6%,提升自然度。
AI 中文摘要
仅解码器的文本到语音(TTS)模型虽然扩展效率高,但在自回归生成过程中,由于文本-语音对齐较弱,容易出现内容幻觉。我们发现,鲁棒性由对齐相关注意力头的锐度之间的非单调关系决定:中等锐度最佳,而过度锐化并不比未对齐的骨干网络更好,甚至鲁棒性更差。基于此,我们提出了AlignDPO,一种后训练方法,通过将轻量级的连接主义时间分类(CTC)对齐项融入直接偏好优化(DPO)中,仅应用于选定的样本,无需架构或推理时的更改,即可达到这一中等锐度区间。在Seed-TTS-Eval英语测试集上,与强DPO基线相比,该方法显著降低了内容幻觉率和词错误率,并将严重内容幻觉率从4.4%降至约0.6%;听力测试进一步表明,其在自然度上优于骨干网络和该基线。因此,对齐最好以中等程度学习并保持,而非最大化或在解码时强制。音频样本可在以下网址获取:https URL。
英文摘要
Decoder-only text-to-speech (TTS) models scale efficiently but remain prone to content hallucinations that arise from weak text-speech alignment during autoregressive generation. We find that robustness is governed by a non-monotone relation to the sharpness of the alignment-bearing attention heads: a moderate degree is best, whereas over-sharpening is no better than the unaligned backbone and even less robust. Guided by this, we present AlignDPO, a post-training method that reaches this moderate regime by folding a lightweight connectionist-temporal-classification (CTC) alignment term into Direct Preference Optimization (DPO), applied only to the chosen samples, with no architectural or inference-time change. On the Seed-TTS-Eval English set, this significantly reduces the content-hallucination and word error rates relative to a strong DPO baseline and lowers the severe content-hallucination rate to ~0.6% (from 4.4%); a listening study further finds it preferred for naturalness over both the backbone and that baseline. Alignment is thus best learned and kept moderate rather than maximized or imposed at decoding. Audio samples are available at https://align-dpo-demo.vercel.app.
Comments6 pages, 4 figures, 4 tables. Accepted to IEEE SLT 2026