发表机构
Stony Brook University; Google DeepMind(纽约州立大学石溪分校; 谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对人类认知中理解与生成的耦合循环,引入自校正耦合马尔可夫跳跃过程框架及$\texttt{CO}_\texttt{2}\texttt{Jump}$采样器,解决掩码扩散模型跨模态矛盾问题,创建多模态语料库,该方法在图像相关任务中性能优异,且性能随去噪步骤数提升。
AI 中文摘要
人类认知不会将理解与生成分开。白板前的教师边说边画,两种模态相互塑造。本文将这种耦合循环引入人工系统。掩码扩散模型(MDMs)很适合此任务,但现有采样器要么交错解码文本和图像,要么在仅共享上一步历史的并行分支中独立更新它们,同一步骤内无法共享另一模态的最新决策,且MDMs无法重新掩码,无法检测和修复跨模态矛盾。我们引入自校正耦合马尔可夫跳跃过程(SC-CMJP)框架,其中一种模态的转移率是另一种模态置信度得分的函数,由跨模态注意力加权。此外,当跨模态证据不利时,重新掩码跳跃会撤回先前的决策。结合SC-CMJP,我们引入了$\texttt{CO}_\texttt{2}\texttt{Jump}$(自校正耦合跳跃),一种用于联合多模态生成的无需训练的单通道采样器。为训练和评估,我们创建并将发布三个大规模联合多模态生成语料库:$\text{JEdit-1M}$、$\text{JMaze-200K}$、$\text{JNono-200K}$,以及匹配的分布内和分布外基准。$\texttt{CO}_\texttt{2}\texttt{Jump}$在图像理解、编辑以及视觉推理(迷宫和数独求解)方面实现了最佳联合性能。采样器的性能随去噪步骤数单调增加,证明跨模态耦合的好处在轨迹上是复合的。项目页面:this https URL
英文摘要
Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws $\textit{together}$, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Masked Diffusion Models (MDMs) are ideally suited to this task, yet existing samplers either decode text and image interleavedly or independently update them in parallel branches that share only previous-step history, but not the other modality's latest decisions $\textit{within}$ the same step; combined with MDMs' inability to remask, cross-modal contradictions are neither detected nor repaired. We introduce $\textbf{Self-Correcting Coupled Markov Jump Processes (SC-CMJP)}$, a framework in which one modality's transition rates are functionals of the other modality's confidence score, as weighted by cross-modal attention. Furthermore, a remasking jump retracts commitments the moment cross-modal evidence turns against them. In conjunction with SC-CMJP, we introduce $\texttt{CO}_\texttt{2}\texttt{Jump}$ (Self-$\underline{\text{CO}}$rrecting $\underline{\text{CO}}$upled $\underline{\text{Jump}}$), a novel training-free single-pass sampler for joint multimodal geneneration. For training and evaluation purposes, we have created and will release three large-scale joint multimodal generation corpora: $\text{JEdit-1M}$, $\text{JMaze-200K}$, $\text{JNono-200K}$, with matching in- and out-of-distribution benchmarks. $\texttt{CO}_\texttt{2}\texttt{Jump}$ achieves best joint performance for image understanding and editing as well as visual reasoning (maze and nonogram solving). The performance of the sampler scales monotonically with the number of denoising steps, evidence that the benefits of cross-modal coupling $\textit{compound}$ across the trajectory. Project page: https://coupled-jump.github.io
CommentsProject page: https://coupled-jump.github.io