发表机构
The Hong Kong University of Science and Technology; Vivix Group Limited; The University of Sydney; Westlake University(香港科技大学; Vivix集团有限公司; 悉尼大学; 西湖大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Salt++提出两阶段后训练框架,通过因果自流和上下文对齐的自回归DMD,解决少步流式多模态生成中的上下文不匹配问题,显著提升视觉与运动质量。
AI 中文摘要
少步流式音频-视频生成需要同时具备因果建模和步骤蒸馏能力,然而标准的训练方案面临两个与上下文相关的挑战。教师强制将干净历史与噪声目标配对,但仅通过速度预测间接监督预测性上下文表示。同时,在因果分布匹配蒸馏(DMD)中直接重用双向评分模型,会造成生成上下文与评分上下文之间的不匹配。我们通过Salt++应对这些挑战,这是一个两阶段后训练框架,包含因果自流(CSF)和上下文对齐的自回归DMD。CSF通过保持噪声目标固定而改变历史来利用上下文信息不对称性:一个噪声混合历史的学生模型将其中间表示与干净历史的指数移动平均教师模型的中间表示对齐。这种自监督信号鼓励学生模型提取语义信息并改善跨模态对齐。上下文对齐的AR DMD在生成器采样、伪评分训练和真实评分评估中共享因果掩码和前缀,以在块条件KL目标下匹配生成分布与参考分布。通过校准的教师引导,它执行干净前缀的少步蒸馏,然后适应生成的历史,而无需切换目标或要求单独的一致性蒸馏。在480p分辨率下,在JavisBench上采用相同的4步因果设置,Salt++相比OmniForcing将视觉质量和运动质量分别提升了57%和45%。一个单独的按尺度后训练阶段将Salt++扩展到4步$1664\ imes960$生成,在七项报告指标中的六项上优于双向LTX-2。项目页面:此https URL
英文摘要
Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Meanwhile, directly reusing bidirectional score models in causal Distribution Matching Distillation (DMD) creates a mismatch between generation and scoring contexts. We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD. CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student aligns its intermediate representations with those of a clean-history exponential-moving-average teacher. This self-supervised signal encourages the student to extract semantic information and improves cross-modal alignment. Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency distillation. At 480p, Salt++ improves visual and motion quality by 57% and 45% over OmniForcing on JavisBench under the same 4-step causal setting. A separate scale-wise post-training stage extends Salt++ to 4-step $1664\times960$ generation, outperforming bidirectional LTX-2 on six of seven reported metrics. Project page: https://xingtongge.github.io/Saltpp
Commentsunder review