arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02208cs.CV

球面编码器2

Sphere Encoder 2

  • University of Maryland(马里兰大学)
  • Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)
  • Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

Kaiyu Yue, Sean McLeish, Ruchit Rawal, Brian Bartoldson, Menglin Jia, Tom Goldstein

AI总结:

Sphere Encoder 2通过解决潜在空间采样间隙和像素级重建损失导致的模糊问题,在保持自编码器速度与简单性的前提下,显著提升图像生成质量。

AI中文摘要:

Sphere Encoder是一种自编码器,通过从高维潜在球体上的随机点解码来生成图像。我们识别出原始公式中限制其生成质量的两个局限性。首先,在编码的潜在空间中,随机点相对于极点更集中于赤道附近,但训练旋转从未到达该区域,留下了一个间隙,限制了一步生成。其次,使用像素级重建损失进行生成训练,鼓励解码器对合理图像进行平均,产生缺乏高频细节的模糊图像。我们提出Sphere Encoder 2来解决这两个局限性,在保持自编码器速度和简单性的同时,大幅提升图像生成质量。模型已在\ref{this https URL}{this http URL}发布。

英文摘要:

Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equator relative to the pole on an encoded latent, but the training rotation never reaches this region, leaving a gap that limits one-step generation. Second, training for generation with pixel-wise reconstruction loss encourages the decoder to average over plausible images, producing blurry images that lack high-frequency details. We present Sphere Encoder 2 to address both limitations, substantially improving image generation quality while maintaining the speed and simplicity of a autoencoder. Models are released at https://github.com/kaiyuyue/sphere2.

补充信息

↑