arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TurboT2VA:基于分数正则化一致性蒸馏的快速大规模文本-视频-音频生成

TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation

Xiaoda Yang, Yuxiang Liu, Kaiwen Zheng, Yuan Liu, Yibo Lai, Shengpeng Ji, Kai Jiang, Jianfei Chen, Shan Yang, Sen Liang, Xiaobin Hu, Shuicheng Yan, Jintao Zhang, Jun Zhu, Zhou Zhao

arXiv 2608.24674首次发表:更新:

发表机构

Zhejiang University; Tianjin University; Tsinghua University; Qingdao University; National University of Singapore(浙江大学; 天津大学; 清华大学; 青岛大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出TurboT2VA框架,通过多模态归一化与渐进式课程方案,在保持音视频生成质量的同时,大幅加速190亿参数联合T2VA模型的推理,在不同分辨率下实现最高54.67倍的生成器加速。

AI 中文摘要

联合文本-视频-音频(T2VA)生成可产出同步的视觉与声学内容,但大型模型的长采样轨迹及异质多模态计算导致推理成本过高。本文提出TurboT2VA,一款用于加速190亿参数联合视频-音频模型的蒸馏与推理框架。大规模T2VA蒸馏面临模态不平衡优化、大规模连续时间一致性训练的难度以及质量-多样性权衡三大挑战。TurboT2VA通过模态归一化及包含离散一致性预热、连续一致性精调、联合一致性-分布匹配的渐进式课程方案解决上述问题:该课程先建立稳定、多样的生成轨迹,再引入分布级精调。在LTX-2数据集上,标准评估分辨率512×768下,四步蒸馏将生成器延迟从50.52秒降至2.51秒,实现20.1倍加速,同时保持优异的视觉质量、音频保真度、多样性及音视频同步性。我们进一步开发架构感知推理栈,结合带保护的W8A8与融合算子、填充文本压缩及模态感知稀疏注意力,同时保留密集跨模态与文本条件路径;在1024×1792高分辨率部署设置下,单NVIDIA H20 GPU上,完整栈将生成器延迟从318.74秒降至5.83秒,实现54.67倍纯生成器加速。推理代码与生成演示可访问该https URL。

英文摘要

Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. We present TurboT2VA, a distillation and inference framework for accelerating a 19B-parameter joint video-audio model. Large-scale T2VA distillation is challenged by modality-imbalanced optimization, the difficulty of continuous-time consistency training at scale, and the quality--diversity trade-off. TurboT2VA addresses these issues with per-modality normalization and a progressive curriculum comprising discrete consistency warm-up, continuous consistency refinement, and joint consistency--distribution matching. The curriculum first establishes a stable, diverse generation trajectory and only then introduces distribution-level refinement. On LTX-2, four-step distillation reduces generator latency from 50.52s to 2.51s at the standard evaluation resolution of 512$\times$768, achieving a 20.1$\times$ speedup while maintaining strong visual quality, audio fidelity, diversity, and video-audio synchronization. We further develop an architecture-aware inference stack that combines guarded W8A8 and fused operators, padded-text compaction, and modality-aware sparse attention while preserving dense cross-modal and text-conditioning paths. Under the high-resolution deployment setting at 1024$\times$1792, the complete stack reduces generator latency from 318.74s to 5.83s on one NVIDIA H20, achieving a 54.67$\times$ generator-only speedup. Inference code and generation demos are available at https://github.com/thu-ml/TurboDiffusion/tree/main/turbot2va.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑