arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38597cs.CV

PixelUMM:无编码器的统一图像与视频理解与生成

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

Cong Wei, Xuanchi Ren, Bryan Chu, Weiming Ren, Huan Ling, Jiahui Huang, Laura Leal-Taixé, Sanja Fidler, Wenhu Chen, Zian Wang, Jay Zhangjie Wu

AI总结:

PixelUMM提出无编码器统一模型,在像素空间直接处理图像与视频,通过空间补丁和时空管状单元及混合Transformer架构,实现理解与生成,实验证明性能竞争力。

AI中文摘要:

统一多模态模型(UMMs)通常依赖分离的视觉表示来进行理解与生成,这增加了视觉上下文的长度,并使与既有视觉-语言预训练流程的集成变得复杂。近期像素空间建模的进展提供了一种无编码器的替代方案,但将该范式从图像扩展到视频并非易事:视频理解与生成采用不同的时间表示,使得统一视觉接口的设计成为一个开放问题。我们提出了PixelUMM,一种直接在像素空间中进行统一图像与视频理解与生成的无编码器模型。PixelUMM将图像表示为空间补丁,将视频表示为时空管状单元,通过单层线性投影将原始像素连接到共享的多模态骨干网络。其Mixture-of-Transformers架构结合了共享注意力与任务特定参数,并将干净像素预测扩展到视频生成,同时支持自回归文本预测和像素空间流匹配。实验表明,PixelUMM在图像与视频理解及生成任务上取得了具有竞争力的性能。我们进一步对关键设计选择进行了实证研究,包括解码器设计以及空间-时间补丁大小,为未来像素空间统一多模态模型提供了见解。

英文摘要:

Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding and generation directly in pixel space. PixelUMM represents images as spatial patches and videos as spatiotemporal tubelets, connecting raw pixels to a shared multimodal backbone through single-layer linear projections. Its Mixture-of-Transformers architecture combines shared attention with task-specific parameters and extends clean-pixel prediction to video generation, jointly supporting autoregressive text prediction and pixel-space flow matching. Experiments show that PixelUMM achieves competitive performance across image and video understanding and generation tasks. We further conduct empirical studies of key design choices, including decoder design and spatial-temporal patch size, providing insights for future pixel-space unified multimodal models.

补充信息

↑