arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

所有模态都是平等的,但视频更平等:缩小联合视频生成中的交叉注意力差距

All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation

Ohad Rahamim, Dvir Samuel, Idan Schwartz, Gal Chechik

arXiv 2609.27901首次发表:更新:

发表机构

Bar-Ilan University; NVIDIA(巴伊兰大学; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出RecCAR正则化方法,通过对齐视频与伴随模态间的交叉注意力,缩小相互对应差距,提升联合视频生成中运动与音频的同步性和质量。

AI 中文摘要

视频是物理事件的丰富表示,捕捉外观、几何、运动和时间的演变。其他模态,如3D身体运动或音频,则编码同一事件的较窄方面。我们发现,联合多模态扩散变换器在跨模态对应上表现出相应的不对称性:伴随模态与视频建立了强对应关系,但通过它们约束视频的相互对应关系仍然明显较弱。我们将两个方向表示为视频令牌上的可比对应分布,并将其分歧定义为相互对应差距。我们引入RecCAR,代表相互交叉注意力正则化,这是一种KL正则化器,利用已建立的视频到模态对应作为固定参考,并将较弱的模态到视频对应向其对齐。在联合视频-运动和视频-音频生成中,RecCAR将人体解剖评分从0.69提高到0.75,并将音频-视频不同步从0.804降低到0.752,同时改善了整体生成质量。

英文摘要

Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Other modalities, such as 3D body motion or audio, encode narrower aspects of the same event. We find that joint multimodal diffusion transformers exhibit a corresponding asymmetry in cross-modal correspondence: companion modalities develop strong correspondences to video, but the reciprocal correspondences through which they constrain video remain substantially weaker. We express both directions as comparable correspondence distributions over video tokens and define their disagreement as the reciprocal correspondence gap. We introduce RecCAR, standing for Reciprocal Cross-modal Attention Regularization, a KL regularizer that uses the well-established video-to-modality correspondence as a fixed reference and aligns the weaker modality-to-video correspondence toward it. Across joint video-motion and video-audio generation, RecCAR improves the Human Anatomy score from 0.69 to 0.75 and reduces audio-video desynchronization from 0.804 to 0.752, while improving overall generation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑