发表机构
Nanyang Technological University; KTH Royal Institute of Technology; Hebei University of Technology(南洋理工大学; 瑞典皇家理工学院; 河北工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出持久性强制(PerF)方法,利用像素空间扩散Transformer中特征专门化的涌现,通过持久特征持续条件化活跃特征,提升图像生成质量,在ImageNet上以更少参数达到接近或更优的FID分数。
AI 中文摘要
像素空间扩散Transformer(DiTs)直接对高维视觉数据进行操作,然而其隐藏表示通常在深度上经历均匀的细化。然而,自然图像本质上是按不同粒度级别组织的。全局结构通常可以紧凑地表示,而局部纹理和精细细节则需要更丰富的表示。受此启发,我们在像素空间DiTs中引入异质细化,为不同的特征组分配不同的细化预算跨深度。因此,出现了一种有序的特征专门化:稀疏细化的特征主要编码全局视觉结构,而更频繁细化的特征逐渐专门化于局部高频细节。我们将这两组分别称为持久特征和活跃特征。基于这种涌现的专门化,我们引入了持久性强制(PerF),它明确利用这种持久-活跃特征组织进行像素空间图像生成。这使得持久特征能够持续条件化活跃细化的特征,允许稳定的全局信息指导更精细视觉细节的持续细化。在生成采样过程中,这种交互进一步诱导出有意义的引导方向,促进连贯的全局结构,并自然地补充了无分类器引导。在ImageNet $256\ imes256$上,PerF-L实现了FID为$1.91$,接近仅用一半参数的JiT-H的$1.86$,而PerF-H在ImageNet $256\ imes256$和$512\ imes512$上进一步实现了FID分别为$1.63$和$1.76$。
英文摘要
Pixel-space diffusion Transformers (DiTs) directly operate on high-dimensional visual data, yet their hidden representations typically undergo uniform refinement across depth. Natural images, however, are inherently organized at different levels of granularity. Global structure can often be represented compactly, whereas local textures and fine details require richer representations. Motivated by this, we introduce heterogeneous refinement in pixel-space DiTs, assigning different feature groups distinct refinement budgets across depth. Consequently, an ordered feature specialization emerges: sparsely refined features predominantly encode global visual structure, whereas more frequently refined features increasingly specialize toward localized, high-frequency details. We refer to these two groups as persistent and active features, respectively. Building on this emergent specialization, we introduce Persistence Forcing (PerF), which explicitly exploits this persistent--active feature organization for pixel-space image generation. This enables persistent features to continuously condition actively refined features, allowing stable global information to guide the ongoing refinement of finer visual details. During generative sampling, this interaction further induces a meaningful guidance direction that promotes coherent global structure and naturally complements classifier-free guidance. On ImageNet $256\times256$, PerF-L achieves FID of $1.91$, approaching $1.86$ of JiT-H with only half the parameters, while PerF-H further achieves FID of $1.63$ and $1.76$ on ImageNet $256\times256$ and $512\times512$, respectively.
CommentsProject page and code: https://chongwang1024.github.io/PerF