arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniGen-AR:通过自回归建模统一视觉生成

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling

Zhipeng Bao, Zhen Zhu, Nupur Kumari, Anurag Bagchi, Yu-Xiong Wang, Pavel Tokmakov, Martial Hebert

arXiv 2607.24157首次发表:更新:

发表机构

Carnegie Mellon University; University of Illinois Urbana-Champaign; Toyota Research Institute(卡内基梅隆大学; 伊利诺伊大学厄巴纳 - 香槟分校; 丰田研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究统一视觉生成问题,提出UniGen-AR框架,将通用多模态语言模型与视觉自回归解码器配对,能为多任务生成图像值输出,相比基于扩散的基线,推理延迟低达19倍,确立视觉自回归建模为统一视觉生成的高效主干。

AI 中文摘要

现代计算机视觉管道仍然碎片化,文本到图像生成、编辑、恢复和经典感知等任务由单独模型处理。我们研究统一视觉生成(UVG),即单个模型通过统一多模态接口产生多样图像值输出。虽然基于扩散的系统因质量和可控性强在UVG中占主导,但迭代采样导致推理延迟大。为此提出UniGen-AR框架,将通用多模态语言模型与高效的下一级视觉自回归解码器配对。该设计保留基于MLLM条件的灵活性,利用VAR模型的采样效率和潜在统一特性。MLLM将自由形式指令和控制信号编码成统一序列,引导VAR解码器为四个类别超过15个任务生成图像值输出。实验上,UniGen-AR推理延迟比基于扩散的基线低达19倍,同时保持或提高输出质量。消融实验还表明VQ-VAE分词器设计对UVG中VAR的可扩展性至关重要。这些结果确立了视觉自回归建模作为统一视觉生成引人注目的高效主干。

英文摘要

Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by separate models. We study Unified Visual Generation (UVG), where a single model produces diverse image-valued outputs through a unified multimodal interface. While diffusion-based systems dominate UVG due to strong quality and controllability, their iterative sampling incurs substantial inference latency, limiting practical deployment. To address these limitations, we propose UniGen-AR, a framework that pairs a general-purpose multi-modal language model (MLLM) with an efficient next-scale visual auto-regressive (VAR) decoder. This design retains the flexibility of MLLM-based conditioning while leveraging the sampling efficiency and latent unification properties of VAR models. In our framework, the MLLM encodes free-form instructions and control signals into a unified sequence, which guides the VAR decoder to generate image-valued outputs for over 15 tasks spanning four families. Empirically, UniGen-AR achieves up to $19 \times$ lower inference latency than diffusion-based baselines while maintaining or improving output quality. Our ablations further reveal that VQ-VAE tokenizer design, particularly codebook size and hierarchy, is a critical factor for VAR scalability in UVG. These results establish visual auto-regressive modeling as a compelling and efficient backbone for unified visual generation. Our project page is at https://zpbao.github.io/projects/unigenar.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑