SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
SVG-T2I: 在视觉基础模型表示空间中扩展文本到图像的潜在扩散模型而不使用变分自编码器
机构 * Department of Automation, Tsinghua University(自动化系,清华大学) ; Kling Team, Kuaishou Technology(快手技术团队)
专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image synthesis(abstract);分类 cs.CV
AI总结 SVG-T2I通过在视觉基础模型表示空间中直接进行文本到图像生成,实现了高质量的图像合成并验证了VFM在生成任务中的能力。
Comments Code Repository: https://github.com/KlingTeam/SVG-T2I; Model Weights: https://huggingface.co/KlingTeam/SVG-T2I