共同原因,而非交叉注意力:阻断音视频生成中的视觉捷径
Intervention, Not Shared Latents: Blocking Visual Shortcuts in Audio-Video Generation
- RIKEN iTHEMS(理化学研究所跨学科理论科学研究所)
- RIKEN AIP(理化学研究所先进智能研究中心)
- Columbia University(哥伦比亚大学)
- South China University of Technology(华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过因果研究证明,音视频生成模型因视觉捷径(从外观而非事件预测声音)而失效,并提出反事实不变性作为识别因果预测器的必要条件,实验验证了干预干扰因素的必要性。
AI中文摘要:
联合音视频生成器训练所用的数据中,事件的外观与其声音之间存在强烈且往往虚假的相关性:特定的材质、纹理或物体外观与特定的声音同时出现。本文针对由此产生的失效模式进行了受控因果研究。我们构建了一个音视频结构因果模型(SCM),其中音频在构造上与视频的干扰外观无关。我们证明,允许音频直接读取视频的模型——通过交叉注意力或共享潜在空间——会学习到一种视觉捷径:它们根据外观而非因果事件来预测声音,并且在测试时当外观-事件相关性被打破时,模型会崩溃,字面意义上合成出错误事件的声音。关键在于,流行的补救措施——将两种模态通过共享的共同原因潜在变量进行路由——并不能解决这个问题:瓶颈结构、无监督的共享/私有分解以及忠实的共享先验模型都会抓住外观代理,并像直接模型一样失败。要阻断捷径,反而需要对干扰因素进行干预。在所述SCM和干预假设下,我们证明了反事实不变性是识别因果预测器的必要且充分条件,并在特征向量SCM、程序化像素视频、带有频谱图音频和预训练骨干的真实图像、移动的真实数字以及条件生成器上验证了该机制。在一个真实的、预训练的视频到音频生成器上,输入干预测试表明,该模型远非对与声音无关的编辑(如对视频重新着色或灰度化会显著改变其生成的声音)具有不变性。
英文摘要:
Joint audio--video (AV) generators are trained on data in which \emph{what an event looks like} and \emph{what it sounds like} are spuriously correlated. We present a \emph{controlled causal study} of the resulting failure mode. In an AV structural causal model where the audio is, by construction, independent of the video's nuisance appearance, models that let audio read video directly---through cross-attention or a shared latent---learn a \emph{visual shortcut}: they predict sound from appearance rather than the causal event and, when the appearance--event correlation is broken at test time, synthesize the wrong event's sound. Crucially, the popular remedy of routing both modalities through a \emph{shared common-cause latent} does \emph{not} fix this---a bottleneck, an unsupervised shared/private factorization, and a faithful shared-prior model all grab the appearance proxy and fail like the direct model. Blocking the shortcut instead requires an \emph{intervention on the nuisance}: under the stated assumptions we prove that counterfactual invariance is necessary and sufficient to identify the causal predictor, and we verify the mechanism from feature-vector SCMs to procedural pixel video, real images with spectrogram audio, moving real digits, and a conditional generator. On a \emph{real, pretrained} V2A generator (MMAudio), an input-intervention test shows the model is far from invariant to sound-irrelevant edits, though a generic-noise control reveals it is broadly input-brittle rather than specifically colour-shortcutting---clean isolation of the shortcut needs the controlled confounds our synthetic studies provide. We characterize \emph{when} the shortcut occurs, compare the objective against supervised counterfactual augmentation, and isolate the \emph{unknown-nuisance} regime---where the intervention cannot be applied---as the central open problem.