AI 中文总结
该研究提出FullDiT模型,通过EMDC和4-CFG等技术解决混合音乐生成器的编解码器接口暴露偏差问题,其系统在多数自动指标上优于商业系统且在人工分析中排名前三。
AI 中文摘要
混合音乐生成器结合了自回归语言模型的长程规划能力与扩散或流基声学渲染器的保真度。然而,渲染器使用干净的、源自目标的编解码器标记进行训练,却在部署时使用不完美的语言模型预测,这造成了编解码器接口暴露偏差。本文未将渲染视为简单的重构任务,而是将其表述为从不完美离散序列出发的全上下文生成任务。我们提出FullDiT,这是一种条件DiT,它将8个帧对齐的RVQ流与独立编码的标题和歌词融合,并在完整的声学潜在序列上使用非因果自注意力。训练期间,误差匹配干扰条件(EMDC)将每个码本的替换率与教师强制的top-1误差率匹配,并从余弦KNN邻域中采样接近错误的标记,且不改变声学目标。推理时,四向无分类器引导(4-CFG)独立缩放编解码器、歌词和标题引导增量。匹配的消融实验显示,在合成损坏下,EMDC使ViSQOL提升0.77,且在与固定语言模型标记的非绑定比较中更受青睐。进一步的消融实验表明,全歌曲上下文和渲染器侧文本条件均带来增益。该完整系统在18个自动指标中的15个上优于5个商业系统,在带人声的音乐人工分析排行榜中位列前三。演示页面可在提供的URL访问。
英文摘要
Hybrid music generators combine the long-range planning of an autoregressive language model with the fidelity of a diffusion- or flow-based acoustic renderer. Yet renderers are trained with clean, target-derived codec tokens but deployed with imperfect language-model predictions, creating codecinterface exposure bias. Rather than treating rendering as a simple reconstruction task,we formulate it as full-context generation from an imperfect discrete plan. We introduce FullDiT, a conditional DiT that fuses eight frame-aligned RVQ streams with independently encoded captions and lyrics and uses non-causal self-attention over the complete acoustic latent sequence. During training, Error-Matched Distractor Conditioning (EMDC) matches per-codebook replacement rates to teacher-forced top-1 error rates and samples near-miss tokens from cosine-KNN neighborhoods without changing the acoustic target. At inference, four-way classifier-free guidance (4-CFG) independently scales codec, lyric, and caption guidance increments. Matched ablations show that EMDC improves ViSQOL by 0.77 under synthetic corruption and is clearly preferred in non-tied comparisons with fixed languagemodel tokens. Further ablations show gains from full-song context and renderer-side text conditioning. The complete system outperforms five commercial systems on 15 of 18 automatic metrics and ranks among the top three on the Artificial Analysis Music with Vocals Leaderboard. The demo page is available at https://selinacloudl.github.io/fulldit-demo/.