发表机构
Zhejiang University; Tianjin University; Tsinghua University; Qingdao University; National University of Singapore(浙江大学; 天津大学; 清华大学; 青岛大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出TurboT2VA框架,通过多模态归一化与渐进式课程方案,在保持音视频生成质量的同时,大幅加速190亿参数联合T2VA模型的推理,在不同分辨率下实现最高54.67倍的生成器加速。
AI 中文摘要
联合文本-视频-音频(T2VA)生成可产出同步的视觉与声学内容,但大型模型的长采样轨迹及异质多模态计算导致推理成本过高。本文提出TurboT2VA,一款用于加速190亿参数联合视频-音频模型的蒸馏与推理框架。大规模T2VA蒸馏面临模态不平衡优化、大规模连续时间一致性训练的难度以及质量-多样性权衡三大挑战。TurboT2VA通过模态归一化及包含离散一致性预热、连续一致性精调、联合一致性-分布匹配的渐进式课程方案解决上述问题:该课程先建立稳定、多样的生成轨迹,再引入分布级精调。在LTX-2数据集上,标准评估分辨率512×768下,四步蒸馏将生成器延迟从50.52秒降至2.51秒,实现20.1倍加速,同时保持优异的视觉质量、音频保真度、多样性及音视频同步性。我们进一步开发架构感知推理栈,结合带保护的W8A8与融合算子、填充文本压缩及模态感知稀疏注意力,同时保留密集跨模态与文本条件路径;在1024×1792高分辨率部署设置下,单NVIDIA H20 GPU上,完整栈将生成器延迟从318.74秒降至5.83秒,实现54.67倍纯生成器加速。推理代码与生成演示可访问该https URL。
英文摘要
Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. We present TurboT2VA, a distillation and inference framework for accelerating a 19B-parameter joint video-audio model. Large-scale T2VA distillation is challenged by modality-imbalanced optimization, the difficulty of continuous-time consistency training at scale, and the quality--diversity trade-off. TurboT2VA addresses these issues with per-modality normalization and a progressive curriculum comprising discrete consistency warm-up, continuous consistency refinement, and joint consistency--distribution matching. The curriculum first establishes a stable, diverse generation trajectory and only then introduces distribution-level refinement. On LTX-2, four-step distillation reduces generator latency from 50.52s to 2.51s at the standard evaluation resolution of 512$\times$768, achieving a 20.1$\times$ speedup while maintaining strong visual quality, audio fidelity, diversity, and video-audio synchronization. We further develop an architecture-aware inference stack that combines guarded W8A8 and fused operators, padded-text compaction, and modality-aware sparse attention while preserving dense cross-modal and text-conditioning paths. Under the high-resolution deployment setting at 1024$\times$1792, the complete stack reduces generator latency from 318.74s to 5.83s on one NVIDIA H20, achieving a 54.67$\times$ generator-only speedup. Inference code and generation demos are available at https://github.com/thu-ml/TurboDiffusion/tree/main/turbot2va.