HiRAE: 具有残差预算的分层表示自编码
HiRAE: Hierarchical Representation Autoencoding with Residual Budgets
AI总结:
HiRAE提出一种分层融合框架,通过残差预算约束编码器各层校正,在保持生成质量的同时显著提升图像重建保真度,并在多个基准上优于现有方法。
AI中文摘要:
预训练的视觉表示支持图像生成,但可能无法完全保留忠实重建所需的细粒度细节。同时,中间编码器层包含互补的视觉细节,但学习融合它们以进行重建可能会产生难以建模的潜在分布。现有的融合方法需要对层选择进行经验性调整,或对融合和解码进行分阶段优化,这增加了配置工作量或训练复杂性。我们引入了HiRAE(分层表示自编码器),它在整个编码器层次上学习一个分层融合框架,以提高重建保真度,同时保持与生成建模的兼容性。HiRAE按深度对编码器层进行分组,并学习对最深表示的残差校正。分组范数上限将这些校正相对于深层锚点进行限制,对较浅的分组采用更紧的预算。我们的HiRAE-24保留了潜在令牌数量和通道维度。在ImageNet-256上,相对于RAEv2,HiRAE-24将重建FID从0.299降低到0.209,同时保持具有竞争力的引导生成质量。对于文本到图像生成,HiRAE-24在GenEval、DPG-Bench和GenAI-Bench上,在监督微调前后均优于RAEv2的对齐性能。在相同的生成器训练和评估协议下,微调后的GenEval从84.86提高到87.70。
英文摘要:
Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding, increasing configuration effort or training complexity. We introduce HiRAE (Hierarchical Representation Autoencoder), which learns a hierarchical fusion framework over the full encoder hierarchy to improve reconstruction fidelity while maintaining compatibility with generative modeling. HiRAE groups encoder layers by depth and learns residual corrections to the deepest representation. Group-wise norm caps bound these corrections relative to the deep anchor, with tighter budgets for shallower groups. Our HiRAE-24 preserves the latent token count and channel dimension. On ImageNet-256, HiRAE-24 reduces reconstruction FID from 0.299 to 0.209 relative to RAEv2 while maintaining competitive guided generation quality. For text-to-image generation, HiRAE-24 improves alignment over RAEv2 on GenEval, DPG-Bench, and GenAI-Bench both before and after supervised fine-tuning. Under the same generator-training and evaluation protocol, post-fine-tuning GenEval increases from 84.86 to 87.70.