发表机构
University of Virginia; Michigan State University; Adobe; Arcade AI(弗吉尼亚大学; 密歇根州立大学; Adobe; Arcade AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PixelDense通过将语义和几何教师模型分离到正交投影流中,实现像素扩散的密集预测对齐,提升生成质量与编辑保真度。
AI 中文摘要
表示对齐(REPA)加速了扩散Transformer的训练,但其对齐目标几乎仅限于语义编码器,如DINOv2和CLIP。近期分析指出,空间结构而非全局语义是对齐效应的载体,然而,为预测该结构而训练的密集预测基础模型作为REPA目标仍被忽视。在像素空间扩散中,SAM2、Depth Anything v2和Metric3D v2各自超越了仅使用DINOv2的GenEval基线,其中两个几何教师模型领先于分割教师模型。然而,将所有四个教师模型进行简单求和,其性能低于最佳单个几何教师模型,因为语义和几何梯度竞争同一个去噪器投影。我们提出了PixelDense,它将DINOv2和SAM2路由到语义投影流,将Depth Anything v2和Metric3D v2路由到几何投影流,并添加了权重空间正交性惩罚,使两个流保持在不相交的子空间中。所有四个教师模型在训练期间冻结,并在推理时丢弃。将PixelDense应用于PixelGen和DeCo,使用单一配方,其在GenEval、DPG-Bench和HPS v2.1上均有所提升,将PixelGen-XXL的GenEval Overall从0.7927提高到0.8093,并超越了所有单教师和未分解的多教师变体。在部分噪声重建中,独立的全景、深度和表面法线探针在COCO和Flickr30K上,在τ=0.5时,PQ增益最高达53.1%,深度AbsRel降低36.0%。从随机初始化开始,PixelDense达到基线峰值GenEval的速度也快了1.23倍。在PIE-Bench上的SDEdit编辑中,PixelDense在每种编辑强度下都保留了更多的源背景和布局,背景PSNR最高提升2.2 dB。
英文摘要
Representation alignment (REPA) accelerates diffusion transformer training, but its alignment targets are almost exclusively semantic encoders such as DINOv2 and CLIP. Recent analysis points to spatial structure, not global semantics, as the carrier of the alignment effect, yet dense-prediction foundation models trained to predict that structure remain overlooked as REPA targets. In pixel-space diffusion, SAM2, Depth Anything v2, and Metric3D v2 each outperform the DINOv2-only GenEval baseline, with the two geometric teachers leading the segmentation teacher. A flat sum of all four teachers, however, lands below the best single geometric teacher, as semantic and geometric gradients compete for one denoiser projection. We introduce PixelDense, which routes DINOv2 and SAM2 through a semantic projection stream, routes Depth Anything v2 and Metric3D v2 through a geometric projection stream, and adds a weight-space orthogonality penalty that keeps the two streams in disjoint subspaces. All four teachers are frozen during training and dropped at inference. Applied to PixelGen and DeCo with a single recipe, PixelDense improves GenEval, DPG-Bench, and HPS v2.1, raises PixelGen-XXL's GenEval Overall from 0.7927 to 0.8093, and beats every single-teacher and unfactored multi-teacher variant. In partial-noise reconstruction, independent panoptic, depth, and surface-normal probes show up to 53.1% PQ gain and 36.0% depth AbsRel reduction at $τ=0.5$ across COCO and Flickr30K. From random initialization, PixelDense also reaches the baseline's peak GenEval 1.23x faster. In SDEdit editing on PIE-Bench, PixelDense keeps more of the source background and layout at every edit strength, raising background PSNR by up to 2.2 dB.
CommentsNeurIPS 2026