arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15395cs.CV

JoLT:用于上下文引导的高分辨率分块生成的联合潜在轨迹

JoLT: Joint Latent Trajectories for Context-Guided High-Resolution Tiled Generation

Mathis Koroglu, Guillaume Jeanneret, Hugo Caselles-Dupré, Matthieu Cord, Arnaud Dapogny

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出JoLT方法,通过联合去噪LR与HR潜在图像生成高分辨率图像,其生成的图像细节丰富、视觉效果佳,优于竞争基线,为艺术创作提供新方向。

中文摘要 AI 辅助

尽管文本到图像生成模型能生成令人印象深刻的结果,但它们难以生成密集细节丰富的高分辨率(HR)图像。现有文献采用从低分辨率到高分辨率的方法解决该问题:首先生成低分辨率(LR)图像,再以该LR图像作为额外线索生成上采样版本。本文提出联合潜在轨迹(JoLT),JoLT生成图像时采用两条流,在每个采样步骤中联合对LR和HR潜在图像去噪:LR潜在控制整体布局,HR潜在控制细节,两条分支相互连接以联合整合信息。我们对该方法进行了广泛验证,证明其相较于竞争基线的优势,生成的图像不仅细节丰富且视觉效果佳,为艺术创作开辟了新途径。

英文摘要

Although text-to-image generative models produce impressive results, they struggle to generate densely detailed, high-resolution (HR) images. Current literature addresses this issue with a low-to-high-resolution approach. First, a low-resolution (LR) image is generated. Then, an upsampled version is generated using the LR image as an additional cue. In this paper, we present Joint Latent Trajectories (JoLT). To generate an image, JoLT uses two streams that jointly denoise LR and HR latent images at each sampling step. The LR latent controls the overall layout, while the HR latent controls the details. We interconnect both branches to jointly integrate their information. We extensively validate our method, demonstrating its advantages over competing baselines. The resulting images are not only richly detailed but also visually pleasing, opening new avenues for artistic creation.

发表机构

  • Obvious Research(奥布弗西斯研究公司)
  • Sorbonne Université(索邦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑