arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22789cs.CV

PixelART:无需潜变量或文本到图像预训练的图像到图层分解

PixelART: Image-to-Layer Decomposition without Latents or Text-to-Image Pretraining

  • Tsinghua University(清华大学)
  • Canva Research(Canva研究院)

机构由 AI 辅助整理,请以论文原文为准。

Zelin Jia, Zhao Zhang, Zhicong Tang, Yuhui Yuan, Shixia Liu

AI总结:

PixelART提出从零训练的像素空间整流流Transformer,无需潜变量或T2I预训练,实现高效图像到RGBA图层分解,参数、延迟和内存大幅降低,性能达到最先进水平。

AI中文摘要:

图像到图层分解将扁平图像转换为可编辑的RGBA图层,从而在设计工作流中实现元素级编辑。现有的基于扩散的系统通常改编大型预训练的文本到图像(T2I)模型,并引入RGBA自编码器或可变图层架构模块。我们重新审视这一设计选择,并质疑图层分解是否真正需要这些重量级组件。我们提出PixelART,一种从零训练的像素空间整流流Transformer,用于图像到图层(I2L)分解。PixelART使用单流多模态扩散Transformer直接对区域RGBA像素块进行去噪,避免了RGBA-VAE、预训练的T2I骨干网络和特定于图层的解码器。我们识别出该任务的一个关键特性:高噪声时间步决定图层分配和粗略图层组织,而低噪声时间步主要细化颜色、透明度、纹理和边界。基于这一观察,我们提出终端增强的时间步采样策略,以增加高噪声图层分配区域的训练覆盖。在400万多层设计模板上训练,PixelART在Design-Multi-Layer-Bench和LICA上实现了最先进的图层分解和合成重建,与最近的Qwen-Image-Layered模型相比,参数减少超过80%,延迟降低98%,内存降低85%。消融实验表明,像素空间x预测、终端增强的时间步采样以及数据/模型缩放至关重要,而T2I预训练对I2L任务带来的收益微乎其微。

英文摘要:

Image-to-layer decomposition converts a flattened image into editable RGBA layers, enabling element-level editing in design workflows. Existing diffusion-based systems typically adapt large pretrained text-to-image (T2I) models and introduce RGBA autoencoders or variable-layer architectural modules. We revisit this design choice and ask whether layer decomposition truly requires these heavyweight components. We introduce PixelART, a pixel-space rectified-flow Transformer trained from scratch for image-to-layer (I2L) decomposition. PixelART directly denoises regional RGBA pixel patches using a single-stream multi-modal diffusion Transformer, avoiding RGBA-VAEs, pretrained T2I backbones, and layer-specific decoders. We identify a key property of the task: high-noise timesteps determine layer assignment and coarse layer organization, while low-noise timesteps mainly refine color, alpha, texture, and boundaries. Based on this observation, we propose a terminal-boosted timestep sampling strategy to increase training coverage in the high-noise layer assignment regime. Trained on 4M multi-layer design templates, PixelART achieves state-of-the-art layer decomposition and composite reconstruction on Design-Multi-Layer-Bench and LICA with over 80% fewer parameters, 98% lower latency, and 85% lower memory than the recent Qwen-Image-Layered model. Ablation experiments show that pixel-space $\mathbf{x}$-prediction, terminal-boosted timestep sampling, and data/model scaling are critical, while T2I pretraining brings marginal benefits to the I2L task.

↑