arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniWorld-Design:从像素生成到层原生设计

UniWorld-Design: From Pixel Generation to Layer-Native Design

Zongjian Li, Zhiyuan Yan, Chenxu Bai, Chen Chen, Haoxiang Sun, Shaodong Wang, Feize Wu, Shenghai Yuan, Bin Lin, Zheyuan Liu, Yuwei Niu, Li Yuan

arXiv 2608.03971首次发表:更新:

AI 中文总结

UniWorld-Design是一种以语义RGBA层为原子单元的图像生成框架,含T2RGBA和I2L两个模型,在Crello基准测试中I2L和T2RGBA均优于现有相关模型。

AI 中文摘要

我们提出UniWorld-Design,这是一个将图像生成从平面像素合成重新定义为结构化视觉合成的框架,以语义RGBA层作为生成、理解和编辑的原子单元。我们的核心见解是:像素定义图像的渲染方式,而层定义图像的创建、理解和编辑方式。正如人类设计师通过层而非原始像素创建和处理视觉内容,UniWorld-Design为多模态生成模型配备了层原生设计空间。UniWorld-Design包含两个模型:文本到RGBA(T2RGBA)模型直接从文本生成独立的RGBA资产;图像到层(I2L)模型以成品图像、全局指令和每层提示为条件,联合生成有序、完整的语义RGBA层。其指令界面支持顶层分解、递归分解和目标提取,使分层成为智能体编辑中可通过指令寻址的操作。由于I2L学习的是完整的语义对象而非可见像素分区,其层在移动或移除时仍保持可用性。在Crello基准测试中,I2L将每层RGB L1误差降低37%,Alpha Soft IoU较Qwen-Image-Layered实现34%的相对提升;单独来看,T2RGBA取得最高CLIP Score,性能优于LayerDiffuse和OmniAlpha。

英文摘要

We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA layers as the atomic units of generation, understanding, and editing. Our key insight is that pixels define how an image is rendered, whereas layers define how an image is created, understood, and edited. Just as human designers create and manipulate visual content through layers rather than raw pixels, UniWorld-Design equips multimodal generative models with a layer-native design space. UniWorld-Design comprises two models. The Text-to-RGBA (T2RGBA) model generates standalone RGBA assets directly from text. The Image-to-Layer (I2L) model conditions on a finished image, a global instruction and per-layer prompts, and jointly produces ordered, complete semantic RGBA layers. Its instruction interface supports top-level decomposition, recursive decomposition and targeted extraction, making layering an instruction-addressable operation for agentic editing. Because I2L learns complete semantic objects rather than visible-pixel partitions, its layers stay usable when moved or removed. On the Crello benchmark, I2L reduces per-layer RGB L1 error by 37% and achieves a 34% relative improvement in Alpha Soft IoU over Qwen-Image-Layered. Separately, T2RGBA achieves the highest CLIP Score, outperforming LayerDiffuse and OmniAlpha.

CommentsProject page: https://rabbitvis.rabbitpre.com/blog

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑