arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从像素出发,无需预训练:单一模型中的联合生成与自监督表示学习

From Pixels, Without Pre-training: Joint Generative and Self-Supervised Representation Learning in One Model

Vicente Balmaseda, Ching-Long Lin, Tianbao Yang

arXiv 2610.05711首次发表:更新:

发表机构

Texas A&M University; University of Iowa(德克萨斯A&M大学; 爱荷华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SCION,在单一模型中联合生成与自监督表示学习,无需标签或预训练模型,通过流时间步条件编码器和嵌入先验实现自条件生成,在ImageNet上以较低FID超越依赖预训练的方法。

AI 中文摘要

强大的图像生成模型通常依赖于类别标签、对齐到冻结的预训练编码器,或建立在单独训练的自编码器之上。尽管这些方法有效,但生成过程依赖于监督或预训练:标签需要人工标注,编码器或自编码器需要针对目标域进行预训练。我们研究在单一模型中联合进行生成和自监督表示学习,从而在没有标签或预训练模型的情况下实现自条件生成。这一目标具有挑战性,因为两种目标函数不匹配:对比学习消耗干净的增强视图,偏好粗粒度、不变性的语义;而流匹配则消耗带噪声的图像,必须保留对比学习所丢弃的精细细节和空间布局。我们提出SCION(基于自监督表示的自条件生成),其核心是一个单一的像素空间编码器,以流时间步和嵌入为条件。对于表示学习,该条件嵌入是一个跨图像共享的学习全局向量,编码器的[CLS]标记产生由对比损失训练的语义表示。对于生成训练,条件嵌入是图像自身的[CLS]表示,而patch标记通过解码器传递以预测图像。为了在推理时无需参考图像进行采样,我们联合学习嵌入上的先验。梯度范数平衡和停止梯度机制使得在一次运行中实现联合优化成为可能。SCION是自监督且自包含的,无需标签或预训练模型。在ImageNet 256x256上,采用JiT-B配方且无表示引导时,SCION达到8.92 FID,超越了无类别条件的iREPA(对齐到预训练DINOv2,FID为46.44)和以其为条件的RCG(FID为14.27)。采用JiT-L时,SCION在无引导下达到5.89 FID,在有表示引导下达到3.47 FID,优于采用ADM配方的RCG(6.24)。

英文摘要

Strong image generation models are conditioned on class labels, aligned to frozen pretrained encoders, or built on separately trained autoencoders. While effective, generation then depends on supervision or pretraining: labels must be annotated, and encoders or autoencoders pretrained for the target domain. We study joint generative and self-supervised representation learning in a single model, enabling self-conditioned generation without labels or pretrained models. This is challenging because the objectives are mismatched: contrastive learning consumes clean augmented views and favors coarse, invariant semantics, while flow matching consumes noisy images and must preserve the fine detail and spatial layout that contrastive learning discards. We propose SCION (Self-conditioned Generation on Self-supervised representation), whose core is a single pixel-space encoder conditioned on the flow timestep and an embedding. For representation learning, this conditioning embedding is a learned global vector shared across images, with the encoder's [CLS] token yielding the semantic representation trained by the contrastive loss. For generative training, the conditioning embedding is the image's own [CLS] representation, while patch tokens pass through a decoder to predict the image. To sample without a reference image at inference, we jointly learn a prior over the embedding. Gradient-norm balancing and stop-gradient mechanisms enable joint optimization in one run. SCION is self-supervised and self-contained, with no labels or pretrained models. On ImageNet 256x256, with the JiT-B recipe and no representation guidance, SCION reaches 8.92 FID, surpassing class-unconditional iREPA, which aligns to pretrained DINOv2 (46.44), and RCG, which conditions on it (14.27). With JiT-L, SCION achieves 5.89 FID without guidance and 3.47 with representation guidance, outperforming RCG with the ADM recipe (6.24).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑