arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15705cs.CV

PixelControl:文本到图像扩散中的细粒度条件保真度

PixelControl: Fine-Grained Condition Fidelity in Text-to-Image Diffusion

Xin Lin, Haodong Li, Zhifei Zhang, Yutong Yang, Haitian Zheng, Juanxi Tian, Zhe Lin, Truong Nguyen

首次发表
浏览论文内容

中文总结 AI 辅助

针对可控文本到图像扩散模型细粒度结构保真度不足的问题,提出PixelControl框架,通过结构感知控制注入与多尺度金字塔循环损失等设计,提升了边界及中/小区域的结构保真度。

中文摘要 AI 辅助

可控文本到图像扩散模型通常能够遵循空间条件的全局布局,但仍会违反物体边界、细轮廓以及中/小尺寸条件区域等细粒度结构。这种限制对于基于VAE的潜在扩散尤其突出,因为空间压缩会削弱高频和小面积条件信号。我们提出PixelControl,一种用于细粒度条件保真度的像素空间可控扩散框架。基于PixelDiT风格的骨干网络,PixelControl避免了潜在瓶颈,并引入两个互补设计:第一,结构感知控制注入,它推导条件结构图并利用其在空间敏感区域周围增强注入的控制残差;第二,多尺度金字塔循环损失,它跨多个分辨率针对条件推导的结构验证生成图像,平衡全局布局一致性与局部边界及细节精度。PixelControl通过具有轻量门控融合的模态特定控制分支支持深度、分割、边缘及其组合。在深度、分割和边缘控制上的实验表明,PixelControl相比现有可控生成方法提升了结构保真度和视觉质量,在边界和中/小尺寸条件区域上增益尤为显著。项目页面可在以下网址获取:this https URL

英文摘要

Controllable text-to-image diffusion models can often follow the global layout of spatial conditions, yet still violate fine-grained structures such as object boundaries, thin contours, and medium/small conditioned regions. This limitation is especially problematic for VAE-based latent diffusion, where spatial compression can weaken high-frequency and low-area condition signals. We propose PixelControl, a pixel-space controllable diffusion framework for fine-grained condition fidelity. Built on a PixelDiT-style backbone, PixelControl avoids the latent bottleneck and introduces two complementary designs. First, Structure-Aware Control Injection derives a condition structure map and uses it to strengthen injected control residuals around spatially sensitive regions. Second, Multi-Scale Pyramid Cycle Loss verifies generated images against condition-derived structures across multiple resolutions, balancing global layout consistency with local boundary and detail accuracy. PixelControl supports depth, segmentation, edge, and their combinations through modality-specific control branches with lightweight gated fusion. Experiments across depth, segmentation, and edge control show that PixelControl improves structural fidelity and visual quality over existing controllable generation methods, with especially strong gains on boundaries and medium/small conditioned regions. The project page can be found at: https://linxin0.github.io/pixelcontrol_homepage/pixelcontrol-site/

发表机构

  • Adobe Research(奥多比研究院)
  • Nanyang Technological University(南洋理工大学)
  • University of California San Diego(加州大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑