音视频模型中的注意力三角
The Attention Triangle in Audio-Video Models
浏览论文内容
中文总结 AI 辅助
本研究针对音视频扩散模型的注意力三角,揭示其双向交互引发的语义泄漏机制,提出注意力衍生信号作为诊断工具并用于推理干预,在保持生成质量的同时提升了跨模态语义对齐效果。
中文摘要 AI 辅助
音视频扩散模型依赖跨模态注意力来协调文本、声音和视觉内容,但该机制会引入微妙且系统性的语义泄漏。本研究通过探查和分析由连接文本、音频和视频流的三条跨注意力边构成的“注意力三角”,探究生成过程中语义信息如何在模态间传递。分析表明,音频-视频边的传递是双向的:音频可影响视频生成,视频也可影响音频生成;该边受模型参数编码的偏差影响,是泄漏的主要来源:当提示与学习到的先验冲突时,跨模态交互可能覆盖预期的条件作用,将语义重定向至视觉规范但不正确的结果。这些效应表明,语义伪影不仅源于注意力超出预期目标的扩散,还源于特定路径上结构化、偏差驱动的交互。基于此视角,我们提取由注意力衍生的信号,这些信号揭示语义如何在模态间分布和落地,并将其作为诊断工具,在受控条件下分析和故意引发泄漏,从而探查跨模态传递的内部动态并分离个体交互的作用。我们进一步利用这些信号指导推理时的干预,以促进更一致的跨模态对齐。大量实验支持我们的分析,证明在保持生成质量的同时提升了语义落地效果。
英文摘要
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes. These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways. Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions. This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment. Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.
发表机构
- Tel Aviv University(特拉维夫大学)
- Simon Fraser University(西蒙菲莎大学)
机构由 AI 辅助整理,请以论文原文为准。