统一多模态模型是否在同一空间中思考?通过跨分支引导的视角
Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering
浏览论文内容
中文总结 AI 辅助
本研究针对统一多模态模型是否共享统一可迁移语义空间的问题,提出跨分支语义引导框架,发现理解分支的引导向量可迁移至生成分支,揭示架构统一不保证语义对齐,该框架可用于探测多模态表示。
中文摘要 AI 辅助
统一多模态模型(UMMs)旨在在单一架构内整合理解与生成能力,但目前仍不清楚这些能力是否共享统一且可迁移的语义空间。该问题极具挑战性,因为两个分支分别处理异构表示(文本令牌与视觉隐变量)及不同的训练目标,难以直接比较。为解决此问题,我们提出了基于干预的框架——跨分支语义引导,该框架从一个分支提取语义方向并将其应用于另一分支。我们表明,从理解分支学习的引导向量可迁移至生成分支,实现可控的图像合成并提升语义保真度;而反向方向的效果始终有限。我们的分析显示,这种不对称性可能与实际的表示不匹配有关:理解分支得到的向量捕获了可迁移的、以对象为中心的语义,而生成分支得到的向量主要编码低级外观特征。我们的结果表明,架构的统一并不能保证语义对齐,并确立跨分支引导作为探测多模态表示的实用工具。
英文摘要
Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet it remains unclear whether these capabilities share a unified and transferable semantic space. This question is fundamentally challenging, as the two branches operate over heterogeneous representations (text tokens vs.\ visual latents) and distinct training objectives, making direct comparison difficult. To address this, we introduce \emph{cross-branch semantic steering}, an intervention-based framework that extracts semantic directions from one branch and applies them to the other. We show that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness. In contrast, the reverse direction consistently shows limited effectiveness. Our analysis suggests that this asymmetry may be related to a practical representational mismatch: understanding-derived vectors capture transferable, object-centric semantics, while generation-derived vectors primarily encode low-level appearance features. Our results reveal that architectural unification does not guarantee semantic alignment, and establish cross-branch steering as a practical tool for probing multimodal representations.
发表机构
- University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。