发表机构
University of California, Los Angeles; Aimakj; National University of Defense Technology(加利福尼亚大学洛杉矶分校; Aimakj; 国防科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出位置解析干扰图和位置感知调制(PAM),在统一多模态模型中按视觉标记位置解析并缓解生成与理解的梯度冲突,实验证明其优于层分离且可互补。
AI 中文摘要
统一多模态模型(UMMs)在共享参数上训练图像理解和自回归图像生成,这两个目标已知会相互干扰。现有的诊断和补救措施在层或专家的分辨率上操作,测量每层的冲突并通过分离参数来解决。我们认为这种分辨率隐藏了一个正交的轴。在UMM中,生成是对视觉标记的光栅序列进行下一个标记预测,这些标记的角色随位置系统性变化,因此生成梯度对理解的干扰程度应取决于其在序列中的来源位置。我们引入了一个位置解析的干扰图,将理解-生成梯度冲突归因于每一层内的视觉标记位置,该图通过一次反向传播计算,其成本为标准反向传播的1.2倍。在Show-o和Janus-Pro上,在控制深度后,位置解释了冲突方差的大部分(在Show-o上,位置的偏η²=0.31,而层的偏η²=0.35;在Janus-Pro上,位置的偏η²=0.15,而层的偏η²=0.30):序列的前四分之一对理解的梯度余弦均值为-0.18,后四分之一为-0.02。这种依赖性在按位置梯度范数归一化后仍然存在,保留了80%的效果大小,且冲突强度与语义内容相关(Spearman ρ=0.64)。基于该图,我们提出了位置感知调制(PAM),它仅在高冲突位置移除生成梯度的反对齐分量,而不改变架构。在匹配的可训练参数预算下,PAM在Show-o上比层分离提高了+21 MME和+2.4 GenEval点,同时在POPE和整体FID上与之匹配;随机位置对照恢复了约31%的增益。基于位置和基于层的分离是互补的自由度,可以结合使用。
英文摘要
Unified multimodal models (UMMs) train image understanding and autoregressive image generation on shared parameters, and the two objectives are known to interfere. Existing diagnoses and remedies operate at the resolution of layers or experts, measuring conflict per layer and resolving it by separating parameters. We argue that this resolution hides an orthogonal axis. Generation in a UMM is next-token prediction over a raster sequence of visual tokens whose roles vary systematically with position, so how strongly a generation gradient interferes with understanding should depend on where in the sequence it originates. We introduce a position-resolved interference map that attributes understanding-generation gradient conflict to visual-token positions within every layer, computed from a single backward pass at $1.2\times$ the cost of a standard backward pass. On Show-o and Janus-Pro, position explains a large share of conflict variance after controlling for depth (partial $η^2=0.31$ vs. $0.35$ for layer on Show-o; $0.15$ vs. $0.30$ on Janus-Pro): the first quarter of the sequence has a mean gradient cosine of $-0.18$ against understanding, the last quarter $-0.02$. The dependence survives per-position gradient-norm normalization, retaining $80%$ of its effect size, and conflict strength tracks semantic content (Spearman $ρ=0.64$). Building on the map, we propose position-aware modulation (PAM), which removes the anti-aligned component of generation gradients only at high-conflict positions without changing the architecture. Under a matched trainable-parameter budget, PAM improves over layer-wise separation by $+21$ MME and $+2.4$ GenEval points on Show-o while matching it on POPE and overall FID; a random-position control recovers about $31%$ of the gain. Position-based and layer-based separation are complementary degrees of freedom and can be combined.
Comments20 pages, 10 figures, 6 tables. Code will be made publicly available upon acceptance