SPAE:用于预训练视觉隐层的频谱引导自编码器
SPAE: Spectrally Guided Autoencoder for Pretrained Visual Latents
浏览论文内容
中文总结 AI 辅助
针对DiT生成隐层与编码器隐层的频谱不匹配问题,提出SPAE框架,通过紧凑瓶颈结构与通道级掩码策略实现隐层适配,在视觉理解、生成质量与重建保真度间取得良好平衡。
中文摘要 AI 辅助
视觉基础模型(Vision Foundation Models, VFMs)生成的隐层具有丰富语义,非常适合视觉理解任务。近期的表征自编码器方法(如RAE)已证实其可生成适用于图像生成的优质隐空间。然而,直接对VFM隐层建模仍存在困难:DiT生成的隐层与编码器隐层存在频谱不匹配问题,尤其在高频分量上。我们的通道级频谱分析进一步发现,这些高频分量在隐层通道中呈分散分布,且与语义信息纠缠,导致DiT难以对该隐空间进行建模。为解决上述挑战,我们提出SPAE,一种用于生成任务的隐层适配框架。具体而言,SPAE采用紧凑瓶颈结构,在提取稳定语义信息的同时抑制高频分量,从而提升DiT生成隐层与编码器隐层的对齐度;此外,我们应用通道级掩码策略,促进瓶颈通道间语义信息与高频细节的解耦。实验表明,SPAE在视觉理解、生成质量与重建保真度之间实现了良好平衡。
英文摘要
Latents from vision foundation models (VFMs) are semantically rich and well suited for visual understanding. Recent representation autoencoder methods such as RAE have shown that they can provide promising latent spaces for image generation. However, VFM latents remain difficult to model directly: DiT-generated latents exhibit spectral mismatch with encoder latents, especially in high-frequency components. Our channel-wise spectral analysis further reveals that these high-frequency components are diffusely distributed across latent channels and entangled with semantic information, making the latent space difficult for DiT to model. To address these challenges, we propose SPAE, latent adaptation framework for generation. Specifically, SPAE employs a compact bottleneck to distill stable semantic information while suppressing high-frequency components, thereby improving the alignment between DiT-generated latents and encoder latents. In addition, we apply a channel-wise masking strategy to promote the decoupling of semantic information and high-frequency details across bottleneck channels. Experiments show that SPAE achieves a favorable balance among visual understanding, generation quality, and reconstruction fidelity.
发表机构
- Alibaba Group(阿里巴巴集团)
- Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。