arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DuSPiT:双分支子补丁像素扩散变换器

DuSPiT: Dual-Branch Sub-Patch Pixel Diffusion Transformer

Yunpeng Bai, Yossi Gandelsman, Michaël Gharbi

arXiv 2607.18510首次发表:更新:

发表机构

The University of Texas at Austin; Reve(德克萨斯大学奥斯汀分校; 瑞夫)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对像素空间扩散中单一表示兼顾全局与细节的问题,提出双分支子补丁像素扩散变换器DuSPiT。通过分离全局与局部处理,利用两分支交互,实现更丰富细节和更好质量效率权衡,优于此前的像素空间扩散变换器。

AI 中文摘要

扩散变换器在图像生成方面性能强大,但大多在压缩潜在空间中运行。像素空间扩散可避免信息丢失,然而现有方法将每个原始图像补丁映射到单个令牌,使单一表示既要处理全局通信又要兼顾细粒度细节。我们提出新架构DuSPiT(双分支子补丁像素变换器)来解决此问题。该模型将全局结构推理与局部外观建模分离,利用紧凑基础分支进行高效全局推理,并行的高容量像素分支(组织成子补丁组)保留详细外观,两分支通过交叉注意力交互。实验结果表明,DuSPiT生成的图像细节更丰富、细粒度结构更强,且在质量-效率权衡上优于先前的像素空间扩散变换器。

英文摘要

Diffusion Transformers achieve strong image generation performance, but most operate in compressed latent spaces. Pixel-space diffusion avoids this information loss, yet existing approaches map each raw image patch to a single token, forcing one representation to handle both global communication and fine-grained details. We address this issue by proposing a new architecture, \textbf{DuSPiT}, a \textbf{Du}al-branch \textbf{S}ub\textbf{P}atch \textbf{Pi}xel \textbf{T}ransformer. This model separates global structural reasoning from local appearance modeling. DuSPiT uses a compact base branch for efficient global reasoning and a parallel, high-capacity pixel branch, organized into subpatch groups, to preserve detailed appearance, with the two branches interacting through cross-attention. Our results show that DuSPiT generates images with richer details and stronger fine-grained structures, while also achieving a better quality--efficiency trade-off than prior pixel-space diffusion transformers.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑