发表机构
Adobe Research(奥多比研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Chimera是一款混合视觉扩散骨干网络,结合KDA、MLA等模块与HeteroP缩放方案,训练出11B参数(2B激活)的模型,在计算效率、视频泛化等方面优于基线,为长上下文扩散架构设计提供基础。
AI 中文摘要
视觉生成日益需要高分辨率图像、长视频及多模态上下文,这使得全注意力机制的二次成本难以承受。我们提出Chimera,一款具备原则性缩放方案的混合视觉扩散骨干网络。Chimera以光栅顺序流处理文本、图像和视频token,无需位置嵌入。它结合了Kimi Delta Attention(KDA)以O(N)复杂度实现长上下文状态跟踪、交错多头潜在注意力(MLA)以实现直接全局交互,以及模态感知的短卷积以捕捉局部时空上下文;同时采用稀疏混合专家(MoE)层在控制激活计算量的同时扩展容量。为对这种异构架构进行缩放,我们提出HeteroP,一种按张量的功能扇入和模型深度跨宽度与深度传递超参数的模块级方案。HeteroP生成一组经一致调优的配置,用于拟合Chinchilla式计算最优规律,涉及激活模型规模、训练token数量及图像-视频数据比例。基于这些规律,我们训练了110亿参数的Chimera,其中20亿为激活参数。实验显示三项结果:第一,以预训练扩散损失衡量,该密集骨干网络的计算效率是匹配的全注意力Wan-2.1 2B基线的1.7倍,完整系统则达到7.3倍;第二,无需针对长度进行微调,Chimera可零样本从5秒训练片段泛化至30秒视频,且最后5秒的FID仅下降6.5%;第三,拟合的规律表明,计算最优的图像预训练会在激活模型规模与训练token数量间近乎均匀分配计算资源,而视频预训练在更高预算下则适度偏向模型规模。这些结果为高效长上下文扩散架构的设计与缩放奠定了基础。
英文摘要
Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with a principled scaling recipe. Chimera processes text, image, and video tokens in one raster-ordered stream without positional embeddings. It combines Kimi Delta Attention (KDA) for long-context state tracking with O(N) complexity, interleaved Multi-head Latent Attention (MLA) for direct global interaction, and modality-aware short convolutions for local spatiotemporal context. Sparse Mixture-of-Experts (MoE) layers expand capacity while controlling activated compute. To scale this heterogeneous architecture, we introduce HeteroP, a module-wise scheme that transfers hyperparameters across width and depth according to each tensor's functional fan-in and model depth. HeteroP yields a consistently tuned family used to fit Chinchilla-style compute-optimal laws for activated model size, training-token count, and image-video data ratio. Guided by these laws, we train an 11B-parameter Chimera with 2B activated parameters. Experiments show three results. First, measured by pretraining diffusion loss, the dense backbone is 1.7x as compute-efficient as a matched full-attention Wan-2.1 2B baseline, while the complete system reaches 7.3x. Second, without length-specific fine-tuning, Chimera extrapolates zero-shot from 5-second training clips to 30-second videos, with only 6.5% FID degradation in the last five seconds. Third, the fitted laws show that compute-optimal image pretraining divides compute nearly evenly between activated model size and training-token count, whereas video pretraining modestly favors model size at higher budgets. These results establish a foundation for designing and scaling efficient long-context diffusion architectures.
Comments40 pages