arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

阿格斯统一模型:迈向紧凑且经济的图像理解与生成统一模型

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation

Weiming Zhuang, Jiabo Huang, Jingtao Li, Zhizhong Li, Chen Chen, Sina Sajadmanesh, Lingjuan Lyu

arXiv 2607.25527首次发表:更新:

发表机构

Sony AI(索尼人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在统一图像理解与生成,提出阿格斯统一模型,利用预训练视觉语言模型,引入混合视觉令牌,分两阶段训练,以最少数据和低成本实现良好性能,达当前最优,可作统一模型开发基线。

AI 中文摘要

将视觉理解和生成统一在一个模型中前景广阔,但因计算和数据需求大以及两种能力所需视觉特征冲突而具有挑战性且成本高昂。为应对这些挑战,我们提出了阿格斯统一模型,这是一个紧凑、有效且统一的多模态模型,对计算和数据需求较低。它有效利用预训练的视觉语言模型提供的多模态先验,引入混合视觉令牌。训练管道包括两个阶段,使用最少的数据量(1560万)和最低成本(约2000美元),在理解和生成方面都取得了良好性能,达到了当前最优水平。我们设想阿格斯统一模型可作为降低统一模型开发障碍的有用基线。

英文摘要

Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data demands and conflicts between the visual features needed for these two capabilities. To address these challenges, we present Argus-Unified, a compact, effective and unified multimodal model built with low demand on computation and data. Instead of aligning modalities from scratch, Argus-Unified effectively leverages pretrained vision-language models (VLMs) that provide strong multimodal priors. Specifically, we introduce hybrid visual tokens that preserve continuous tokens for understanding while learning discrete tokens for generation from a frozen unified vision encoder. Our training pipeline includes two stages: the first stage learns a quantizer and image decoder on top of the frozen vision encoder, the second stage trains the LLM initialized from a pretrained VLM for the unified multimodal modeling. Using by far the least amount of data (15.6M) and the lowest cost (~$2,000), we demonstrate that unified multimodal models can be trained economically while achieving strong performance in both understanding and generation. Notably, our model attains state-of-the-art multimodal understanding on GQA, POPE, and VQAv2, and competitive generation quality compared to models with dedicated vision encoders (e.g., Janus, Janus-Pro), all at ~10x lower cost and with ~5x less data. We envision Argus-Unified as a useful baseline that lowers the development barrier for unified models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑