发表机构
Computer Vision Center; Universitat Autònoma de Barcelona; City University of Hong Kong (Dongguan); Universitat de València(计算机视觉中心; 巴塞罗那自治大学; 香港城市大学(东莞); 瓦伦西亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文发现文本到图像模型的VAE潜在空间共享与亮度和对立色对齐的颜色子空间,并据此提出ColorTuning、饱和度控制和颜色迁移三种应用,其中ColorTuning在精确颜色生成上达到最先进性能。
AI 中文摘要
变分自编码器(VAEs)是现代文本到图像模型的关键组成部分,这些模型在其潜在空间内生成图像。已知VAEs能够解耦数据中的主要变化因素,而颜色被认为是自然图像中这些因素中结构最明显的之一:去相关化产生一个亮度轴和两个对立色轴。因此,颜色应预期在VAE潜在空间中作为一个独立因素出现。然而,这些潜在空间如何表示颜色在很大程度上仍未被探索。在本工作中,我们展示了文本到图像模型的VAEs共享一个与亮度和对立色对齐的颜色子空间。通过编码器的线性近似和有针对性的潜在空间引导,我们在广泛的VAEs中一致地发现了这一子空间,从SD1.5到FLUX.2和Z-Image。基于这一表征,我们提出了三个应用:ColorTuning,在GenColorBench的细粒度CSS3/X11系统上实现了精确数值颜色生成的最先进性能;饱和度控制,用于调整全局色彩强度;以及颜色迁移,用于将调色板更改为匹配参考。代码和模型可在该https URL公开获取。
英文摘要
Variational autoencoders (VAEs) are a key part of modern text-to-image models, which generate images within their latent space. VAEs are known to disentangle the main factors of variation in the data, and color is known to be one of the most structured of these in natural images: decorrelating it yields one luminance axis and two opponent-color axes. Color should therefore be expected to emerge as a distinct factor in the VAE latent space. Yet how these latent spaces represent color remains largely unexplored. In this work, we show that the VAEs of text-to-image models share a color subspace aligned with brightness and opponent-colors. Through a linear approximation of the encoder and targeted latent steering, we find this subspace consistently across a broad range of VAEs, from SD1.5 to FLUX.2 and Z-Image. Building on this characterization, we propose three applications: ColorTuning, which achieves state-of-the-art in precise numerical color generation on the fine-grained CSS3/X11 system of GenColorBench, saturation control, to adjust the global chromatic intensity, and color transfer, to change the palette to match a reference. The code and models are publicly available at https://julian075.github.io/Color_Subspace/