新概念必须进入之处:统一多模态模型中的入口门跨任务可用性
Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models
浏览论文内容
中文总结 AI 辅助
该研究通过分离统一多模态模型的理解与生成任务方向,发现跨任务可用性取决于概念绑定的入口层,提出的对齐目标可在极低损失下实现概念跨任务迁移。
中文摘要 AI 辅助
统一多模态模型(UMMs)的研究动机是希望理解与生成任务能相互促进,但大量控制消融实验发现,加入生成目标后,理解性能并未提升。联合训练研究无法解决这一争议:在存在重叠监督的情况下,无法将性能提升归因于模型架构而非数据。为进一步探究UMMs中两个任务方向的关系,本文通过构造将二者分离:将一个新颖的视觉实体(渲染的3D资产与伪词配对,且该伪词经筛选后不会出现在冻结模型的行为中)通过恰好一个任务方向绑定,随后测量未训练方向的表现。研究发现,两个方向的通道均真实存在,但性质不同:生成训练会让模型获得仅能在候选中匹配的名称,而理解训练会让模型获得还能生成的名称。跨任务可用性由绑定进入共享计算的位置决定。对齐探针可预测36种配置下的迁移效果(Spearman相关系数ρ=+0.68)。该目标的对齐项在所有权重冻结的激活上以闭式形式最大化,当在28层的第7层注入概念时,可使该概念被模型识别,而在第14层及以上注入时,模型表现与基础模型无差异;而相同编辑的基于权重的版本则在第10-14层达到峰值。在对四个模型的观测序列中,仅当理解通路为语义视觉编码器时,该窗口才会出现,这表明统一权重并不足够:两个方向必须在入口点共享语义格式。利用该规则,中间层对齐目标可使模型获得该概念,仅导致模型通用文本到图像能力的相对损失为0.1%,而标准生成路线的相对损失为41%。本文代码位于该https URL。
英文摘要
Unified multimodal models (UMMs) are motivated by the hope that understanding and generation reinforce each other but controlled ablations repeatedly find that adding a generation objective leaves understanding flat. Joint-training studies cannot settle the disagreement: with overlapping supervision, a gain cannot be attributed to the architecture rather than the data. To further investigate the relationship between the two directions in UMMs, we separate them by construction. A novel visual entity, a rendered 3D asset paired with a pseudo-word screened for absence from the frozen model's behavior, is bound through exactly one task direction, and the untrained direction is then measured. We find that the channel is real in both directions, but the directions differ in kind: generation training installs a name the model can only match among candidates; understanding training installs one it can also produce. What governs cross-task usability is where the binding enters the shared computation. An alignment probe predicts export across 36 configurations (Spearman $ρ= +0.68$). That objective's alignment term, maximized in closed form over activations with every weight frozen, makes a concept drawable when injected at layer 7 of 28 and is indistinguishable from the base model from layer 14 on, while the weight-based version of the same edit peaks at layers 10-14. In an observational series of four models, this window appears only where the understanding pathway is a semantic vision encoder, suggesting that unified weights are not enough: the two directions must share a semantic format at the entry point. Exploiting the rule, a mid-stack alignment objective acquires the concept for a $0.1\%$ relative loss of the model's general text-to-image ability, against $41\%$ for the standard generative route. Our code is at https://github.com/Zane-ZYQiu/entry-point-umm.
发表机构
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- Columbia University(哥伦比亚大学)
- CUHK(香港中文大学)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。