FuseReg:正则化层融合缓解表示自编码器中的重建-生成差距
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
浏览论文内容
中文总结 AI 辅助
FuseReg通过随机子集训练正则化层融合,缓解表示自编码器的重建-生成差距,仅替换解码器即可降低27%的gFID,联合正则化可降低29%。
中文摘要 AI 辅助
表示自编码器(RAE)复用预训练视觉编码器的特征作为重建和扩散潜变量,将强大的视觉表示整合到图像生成中。然而,RAE仍需决定哪些编码器层构成生成器和像素解码器的共享潜空间。这一选择涉及权衡:较浅的层往往能更好地保留精细像素细节,而较深的层往往能产生更好的生成指标。因此,固定的启发式层融合将两个受益于不同信息的阶段耦合在一起。我们提出FuseReg,用对编码器层随机子集的训练替代启发式特征选择。我们从理论上分析了其底层机制:子集采样显式惩罚了对跨层不一致性的敏感度。在ImageNet-256上使用DINOv3-L,单个FuseReg解码器无需重新训练即可从完整、稀疏和单层融合中重建,其PSNR高于针对固定融合特化的解码器。这种灵活性也有利于生成:仅替换解码器即可在保持RAEv2 DiT-XL生成器不变的情况下,将无引导gFID降低27%。同样的正则化原则可扩展到扩散训练,对两个阶段进行联合正则化可将DiT-Base上的无引导gFID降低29%。这些结果表明,针对层融合鲁棒性训练下游模型可在不修改预训练编码器的情况下缩小重建-生成差距。
英文摘要
Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off: shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling preserves the full-layer latent mean in expectation while explicitly penalizing sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. Applying FuseReg to both stages also reduces unguided gFID by 29% on DiT-Base. The reconstruction and generation benefits also extend to other encoder families. FuseReg narrows the reconstruction-generation gap without additional training cost or architectural changes.
发表机构
- USC PSI Lab(南加州大学PSI实验室)
- Brown University(布朗大学)
- Rice University(莱斯大学)
- University of Aberdeen(阿伯丁大学)
- University of Notre Dame(圣母大学)
- University of Maryland, College Park(马里兰大学帕克分校)
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。