发表机构
University of Geneva; Meta(日内瓦大学; Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出PoolDINO,一种学习型仿射池化算子,通过合并令牌实现高效扩散生成,在ImageNet-256上实现4倍压缩保持质量,16倍压缩提升效率,吞吐量最高提升9倍。
AI 中文摘要
表示自编码器(RAEs)从预训练的视觉特征生成图像,但其密集的令牌网格使得生成建模成本高昂。受局部特征相关性的启发,我们引入了PoolDINO,一种学习得到的仿射池化算子,用于合并相邻令牌。将池化算子与RGB解码器联合训练,保留了标准的两阶段RAE流程,无需单独的特征自编码器。在ImageNet-256上,4倍令牌压缩在内部引导下保持了相当的生成质量,而16倍压缩则以一定的质量换取更高的效率。在固定的100步采样预算下,相对于未池化的基线,潜在空间采样吞吐量分别提高了3.7倍和9.0倍。分类和密集预测评估表明,相当的引导生成质量可以与在其他任务上较弱的性能共存。
英文摘要
Representation Autoencoders (RAEs) generate images from pre-trained visual fea- tures, but their dense token grids make generative modeling expensive. Motivated by local feature correlations, we introduce PoolDINO, a learned affine pooling operator that merges neighboring tokens. Training the pooling operator jointly with the RGB decoder preserves the standard two-stage RAE procedure without a separate feature auto-encoder. On ImageNet-256, 4x token compression retains comparable generation quality under internal guidance, while 16x compression trades some quality for greater efficiency. At a fixed budget of 100 sampling steps, latent-sampling throughput increases by 3.7x and 9.0x, respectively, relative to the unpooled baseline. Classification and dense prediction evaluations show that comparable guided generation quality can coexist with weaker performance on other tasks.