发表机构
Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 UniSpace,通过 Patch Reparameterization 改进语义 ViT,构建 80 亿参数的 Transformer 专家混合模型,实现同一视觉空间内的理解、生成与编辑,提升重建效果,支持文本到图像生成和指令式图像编辑。
AI 中文摘要
语义视觉编码器已成为多模态理解和图像生成中语义条件设置的核心视觉接口,但它们的最终 token 会丢弃细粒度视觉细节,导致像素重建效果差,限制了其在图像生成和编辑等对重建敏感的任务中的应用。本研究探讨能否在基于预训练语义 ViT 构建的单一视觉表示空间中对理解、生成和编辑进行建模。研究发现,语义 ViT 的冻结 Transformer 模块并非本身无法保留视觉细节,而是原始的 patch 参数化会将表示推向语义抽象,导致最终 token 难以恢复细粒度信息。基于该观察,本文提出 Patch Reparameterization(patch 重参数化)方法,在保留原有语义通路的同时,为相同的冻结 ViT 模块添加具备重建感知能力的 patch 嵌入,以提供细粒度视觉信息。由此得到的统一表示既保留了多模态理解能力,又能实现高保真图像重建,且具备良好的重建-生成权衡效果。本文进一步将该表示扩展为 UniSpace,这是一个拥有 80 亿参数的 Transformer 专家混合模型,无需单独的 VAE 通路,可在同一视觉空间中完成理解、生成和编辑任务。系统级评估表明,该模型可实现实用的文本到图像生成和基于指令的图像编辑,证明重参数化后的预训练 ViT 可作为统一视觉接口用于可扩展多模态建模。
英文摘要
Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, their final tokens discard fine-grained visual details, leading to poor pixel reconstruction and limiting their use in reconstruction-sensitive tasks such as image generation and editing. In this work, we ask whether understanding, generation, and editing can be modeled in a single visual representation space built from a pretrained semantic ViT. We show that the frozen Transformer blocks of a semantic ViT are not intrinsically unable to preserve visual details. Instead, the original patch parameterization drives the representation toward semantic abstraction, making fine-grained information difficult to recover from the final tokens. Based on this observation, we introduce \emph{Patch Reparameterization}, which preserves the original semantic pathway while adding a reconstruction-aware patch embedding that provides fine-grained visual information to the same frozen ViT blocks. The resulting unified representation preserves multimodal understanding while enabling high-fidelity image reconstruction and a favorable reconstruction--generation trade-off. We further scale this representation into \emph{UniSpace}, an 8B Mixture-of-Transformer-Experts model that performs understanding, generation, and editing in the same visual space without a separate VAE pathway. System-level evaluations demonstrate practical text-to-image generation and instruction-based image editing, showing that a reparameterized pretrained ViT can serve as a unified visual interface for scalable multimodal modeling.