AI 中文总结
研究统一多模态模型中理解与生成的对齐差距,提出STBridge共享目标对齐框架,通过共同目标状态连接二者,采用对齐然后优化策略,经实验验证该框架能缩小模型描述与生成间差距,有效弥合理解与生成。
AI 中文摘要
统一多模态模型旨在在单一架构中集成视觉理解和生成,但仅架构统一并不能确保语义一致性。模型可能正确描述目标但生成不一致的编辑,这暴露了理解与生成的对齐差距。在图像编辑中研究此差距,通过比较目标字幕和编辑后的图像来测试输出是否收敛。分析表明现有模型对齐性弱。提出STBridge框架,通过共同目标状态连接理解与生成,采用对齐然后优化策略,在多个基准测试中持续改进,缩小了描述与生成之间的差距。
英文摘要
Unified multimodal models (UMMs) aim to integrate visual understanding and generation within a single architecture, but architectural unification alone does not ensure semantic consistency. A model may describe the intended target correctly while generating an inconsistent edit. This exposes an understanding-generation alignment gap: linguistic and visual outputs live in different spaces, yet should be governed by the same target semantics. We study this gap in image editing, where an instruction defines a target state that can be both described and visually realized. Given a source image and an edit instruction, we compare a UMM's target caption with its edited image to test whether the two outputs converge on the same result. Our analysis shows that existing UMMs remain weakly aligned, especially for fine-grained entities, attributes, spatial relations, and local details, indicating that semantic unification is not achieved by architecture alone. To bridge this gap, we propose STBridge, a shared-target alignment framework that connects understanding and generation through a common target state. Here the target caption expresses the desired visual result, while the edited image realizes it visually, replacing separate task-specific paths with a shared information flow from target expression to target realization. STBridge follows an align-then-optimize strategy: supervised fine-tuning first establishes the shared-target channel, and sequential reinforcement learning further refines target-centered coordination. Across visual understanding, image generation, and image editing benchmarks, STBridge consistently improves over the initialization model. Alignment analysis confirms that STBridge narrows the gap between what the model describes and what it generates, demonstrating shared-target alignment as an effective post-training strategy for bridging understanding and generation in UMMs.