利用内在主体感知注意力实现可控多主体视频生成
Harnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video Generation
- Alibaba Group(阿里巴巴集团)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对多主体视频生成中的保真度不可控和语义漂移问题,提出DIAL框架,利用扩散变换器内在注意力图实现可控生成,在OpenS2V-Eval上显著提升身份一致性。
中文摘要 AI 辅助
多主体视频生成面临两个关键挑战:不可控的保真度强度和潜在语义漂移。我们通过分析扩散变换器(DiTs)的内部机制来解决这些问题。我们发现,某些注意力块自然地形成了一种内在空间定位图(ISGM),能够精确定位参考主体。基于这一洞察,我们提出了双阶段内在注意力利用(DIAL)框架,该框架在训练和推理阶段均使用这些内部信号。在低噪声阶段,我们使用ISGM引导注意力机制,从而在无需重新训练的情况下,在推理期间实现对保真度强度的精确控制。在高噪声阶段,我们利用这些相同的图自动构建偏好对,用于强化学习(RL),且无需额外成本。这一强化学习过程有效地将模型的注意力锚定到参考主体上,并缓解了语义漂移。大量实验表明,DIAL在OpenS2V-Eval基准上显著优于基线模型,持续提升了身份一致性,并实现了可控的保真度强度。
英文摘要
Multi-subject video generation faces two key challenges: uncontrollable fidelity strength and potential semantic drift. We address these by analyzing the internal mechanisms of Diffusion Transformers (DiTs). We found that certain attention blocks naturally form an Intrinsic Spatial Grounding Map (ISGM) that precisely locates reference subjects. Building on this insight, we propose Dual-phase Intrinsic Attention Leveraging (DIAL), a framework that uses these internal signals for both training and inference. In low-noise stages, we use ISGM to guide the attention mechanism, allowing precise control over fidelity strength during inference without retraining. In high-noise stages, we use these same maps to automatically build preference pairs at no additional cost for Reinforcement Learning (RL). This RL procedure effectively anchors the model's attention to reference subjects and mitigates semantic drift. Extensive experiments show that DIAL significantly outperforms baseline models on the OpenS2V-Eval benchmark, consistently improving identity consistency and enabling controllable fidelity strength.