arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.31106cs.CVcs.SD

DreamX-Creator:实现2K分辨率原生音视频生成的民主化

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang, Shen Zhang, Bingze Song, Jiachen Lei, Ruimin Lin, Jiahong Wu, Xiangxiang Chu

首次发表
浏览论文内容

中文总结 AI 辅助

DreamX-Creator 1.0是基于7B生成器的紧凑原生联合音视频生成系统,通过多阶段训练与2K细化流水线实现2K分辨率同步音视频生成,性能达开源先进水平,旨在推动该领域研究民主化。

中文摘要 AI 辅助

近期的视频生成器往往会忽略音频,或将音频在单独阶段合成,这限制了视觉动态与声学事件之间的双向建模。我们提出了DreamX-Creator 1.0,这是一个以7B生成器为核心的紧凑原生联合音视频生成系统。在首帧和文本提示的条件下,该生成器对模态专用的音频和视频流进行联合去噪。这些流在网络前半部分独立处理,后半部分通过门控跨模态注意力(Gated Cross-Modal Attention)耦合,该注意力的令牌级和头级输出门会调制每个活跃跨模态注意力头的输出。统一音视频数据系统构建并过滤时间连贯的片段,生成结构化多模态注释,并将片段组织为面向能力的数据池。渐进式联合训练包含两个音视频预训练阶段,随后是高质量微调。音视频强化学习进一步通过模态感知多模态反馈对生成器进行后训练,该反馈将视频、音频和跨模态反馈路由到相应的流。对于高分辨率输出,我们的自回归1步2K细化(Autoregressive 1-Step 2K Refinement)流水线将双向多步教师模型适配为自回归多步细化器,并将其蒸馏为学生模型,该学生模型每个时间块仅需一次去噪评估。总体而言,DreamX-Creator 1.0实现了原生、同步的音视频生成,性能可与最先进的开源系统相媲美。通过发布我们紧凑的7B生成器和2K细化器,我们力求实现原生音视频生成的民主化,并为未来统一音视频生成建模的研究提供可访问的基础。

英文摘要

Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generator jointly denoises modality-specialized audio and video streams. The streams are processed independently in the first half of the network and coupled in the latter half through Gated Cross-Modal Attention, whose token- and head-wise output gates modulate each active cross-modal attention-head output. A unified Audio-Video Data System constructs and filters temporally coherent clips, produces structured multimodal annotations, and organizes clips into capability-oriented data pools. Progressive Joint Training comprises two audio-video pre-training stages followed by High-Quality Finetuning. Audio-Video Reinforcement Learning further post-trains the generator with Modality-Aware Multimodal Feedback that routes video-, audio-, and cross-modal feedback to the corresponding streams. For high-resolution output, our Autoregressive 1-Step 2K Refinement pipeline adapts a bidirectional multi-step teacher into an autoregressive multi-step refiner and distills it into a student requiring one denoising evaluation per temporal chunk. Overall, DreamX-Creator 1.0 achieves native, synchronized audio-video generation with performance competitive with state-of-the-art open-source systems. By releasing our compact 7B generator and 2K Refiner, we seek to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.

发表机构

  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

↑