发表机构
UC San Diego; Shanghai Jiao Tong University; Aether AI(加州大学圣迭戈分校; 上海交通大学; Aether AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出残差流负担机制,解释扩散Transformer中预测目标与架构如何共同影响表示学习,并引入空间索引超连接(SiHC)在ImageNet 256×256上实现FID 1.71。
AI 中文摘要
在基于扩散的生成中,神经网络可以被训练为从带噪声的输入预测干净数据、噪声或速度。这些预测目标可以相互转换并描述相同的生成过程,然而在大型像素块上操作的普通扩散Transformer在干净预测上成功,在噪声或速度预测上失败。我们认为这种不对称性源于噪声目标要求残差流在深度方向上保留依赖于噪声的输入变化以供最终读出,迫使后续层在噪声表示上计算。频谱集中的干净目标施加了较轻的需求,留下更大的自由度来组织隐藏表示以供后续计算。我们将这种保留需求称为*残差流负担*,并展示它如何塑造扩散Transformer中的表示学习。受控实验表明,可利用的结构是补丁空间中的频谱集中性,而持久残差状态的带宽是噪声预测的关键资源。我们进一步表明,这一解释与最近的解耦像素空间架构一致,这些架构的多样化设计都减少了主路径上的残差流负担。为了从互补方向检验这一理解,我们直接扩展和重组残差流带宽,引入了空间索引超连接(SiHC),在ImageNet $256^2$上达到FID 1.71。综合这些结果,我们确定残差流负担是预测目标和架构共同塑造扩散Transformer中表示学习的一种机制。
英文摘要
In diffusion-based generation, a neural network can be trained to predict the clean data, the noise, or the velocity from a noisy input. These prediction targets are interconvertible and describe the same generative process, yet plain Diffusion Transformers operating on large pixel patches succeed with clean prediction and fail with noise or velocity prediction. We argue that this asymmetry arises because noisy targets require the residual stream to preserve noise-dependent input variation through depth for the final readout, forcing subsequent layers to compute on noisy representations. A spectrally concentrated clean target imposes a lighter demand, leaving greater freedom to organize hidden representations for subsequent computation. We call this preservation requirement *residual-stream burden* and show how it shapes representation learning in Diffusion Transformers. Controlled experiments indicate that the exploitable structure is spectral concentration in patch space and that the bandwidth of the persistent residual state is a key resource for noisy prediction. We further show that this account is consistent with recent decoupled pixel-space architectures, whose diverse designs all reduce the residual-stream burden on the main pathway. To examine this understanding from a complementary direction, we expand and reorganize the residual-stream bandwidth directly, introducing Spatially Indexed Hyper-Connections (SiHC) that reach FID 1.71 on ImageNet $256^2$. Together, these results identify residual-stream burden as a mechanism through which prediction targets and architecture jointly shape representation learning in Diffusion Transformers.
CommentsUnder review. Repo release: https://github.com/tongtongliang/residual-stream-burden