arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09450cs.CVcs.AI

Iris-3B:超越潜在空间,走向像素空间扩散训练、转换与微调

Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

Hanqiu Li Cai, Chema Garabito

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Iris-3B,一个30亿参数的像素空间文本到图像模型,通过课程预训练和潜在模型转换,验证像素空间先验在下游任务中无显著优势,但达到与潜在模型竞争的质量。

中文摘要 AI 辅助

像素空间扩散模型避免了潜在模型的有损VAE,这暗示了在细粒度细节重要的下游任务上具有优势。我们沿着两条通往像素空间骨干网络的路径测试了这一说法。我们从头预训练了Iris-3B,一个30亿参数的像素空间文本到图像Transformer,通过$256\ o512\ o1024$的课程,首先在$256^2$分辨率下消融了预测目标和表示对齐,以决定扩展什么。我们还将预训练的潜在模型FLUX.2 Klein base 4B转换为像素空间。我们对这两个家族进行微调,用于单目深度估计和图像恢复/超分辨率。我们发现使用像素空间生成先验没有显著改进。使用一种匹配的直接回归方法进行深度微调,Iris-3B与潜在FLUX.2 Klein持平,而转换后的像素FLUX.2 Klein则落后于它;在$4\ imes$ DIV2K恢复任务上,两个像素模型均未击败潜在FLUX.2 Klein微调,转换后的模型略微落后。我们记录了配方、失败模式和这一负面结果背后的剩余混淆因素。尽管如此,Iris-3B表明,使用PixelDiT的像素Transformer(PiT)头进行像素空间预训练可扩展到30亿参数,并达到与潜在模型竞争的文本到图像质量,在$1024^2$分辨率下,在OneIG上匹配Qwen-Image的官方评估器。我们发布其权重和训练代码,希望它们有助于为像素空间生成的进一步工作铺平道路。

英文摘要

Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters. We test this claim along both routes to a pixel-space backbone. We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a $256\to512\to1024$ curriculum, after first ablating the prediction target and representation alignment at $256^2$ to decide what to scale. We also convert a pretrained latent model, FLUX.2 Klein base 4B, to pixel space. We fine-tune both families for monocular depth estimation and for image restoration/super-resolution. We find no significant improvement from using a pixel-space generative prior. Fine-tuned for depth with one matched direct-regression recipe, Iris-3B is level with the latent FLUX.2 Klein and the converted pixel FLUX.2 Klein falls behind it, and on $4\times$ DIV2K restoration neither pixel model beats a latent FLUX.2 Klein fine-tune, the converted one trailing it slightly. We document the recipes, the failure modes and the remaining confounds behind this negative result. Nevertheless, Iris-3B shows that pixel-space pretraining with the pixel-transformer (PiT) head of PixelDiT scales to 3B parameters and to text-to-image quality competitive with latent models, matching Qwen-Image on OneIG under the official evaluators at $1024^2$. We release its weights and training code in the hope that they help pave the way for further work on pixel-space generation.

发表机构

  • SperidLabs

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑