arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LiFT:循环流变换器

LiFT: Loop Flow Transformers

Mohammad Mahdi Derakhshani, Pedro M. P. Curvo, Gertjan J. Burghouts, Jan-Willem van de Meent, Cees G. M. Snoek

arXiv 2610.05538首次发表:更新:

发表机构

VISLab; University of Amsterdam; TNO, Intelligent Imaging; AMLab(VISLab; 阿姆斯特丹大学; 荷兰国家应用科学研究院,智能成像; AMLab)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LiFT 通过重复应用共享 DiT 核心并采用直线路径回归目标,使模型可循环超过训练深度,在不增加参数的情况下提升生成质量,在 ImageNet 上以更少参数和计算量显著降低 FID。

AI 中文摘要

我们提出了循环流变换器(LiFT),这是一类循环生成模型家族,通过重复应用共享的扩散变换器(DiT)核心来扩展计算量,仅需对标准架构进行少量改动。LiFT 并非要求每个循环步骤都给出最终预测,而是用单一回归目标训练每一步:该目标是从模型初始估计到流匹配目标的一条直线路径上的一个点。由于我们通过连续的深度坐标来索引这些目标,训练后的模型可以在无需重新训练、提前退出或其他修改的情况下,循环深度远超其训练深度。在我们的实验中,这些更长的循环展开改善了生成效果,因此推理计算可以在不增加参数的情况下增长。在 ImageNet 256x256 分辨率下,LiFT-L/2 的 FID 比我们密集的 DiT-XL/2 基线低 3.34 点,同时使用的参数减少约 60%,训练 FLOPs 减少 32%,推理 FLOPs 减少 52%。

英文摘要

We introduce Loop Flow Transformers (LiFT), a family of looped generative models that scales computation by repeatedly applying a shared Diffusion Transformer (DiT) core, with only light changes to the standard architecture. Rather than asking every recurrent step for the final prediction, LiFT trains each step with a single regression target: a point on a straight path from the model's initial estimate to the flow-matching target. Because we index these targets by a continuous depth coordinate, a trained model can loop far beyond its training depth with no retraining, early exits, or other modifications. In our experiments, these longer rollouts improve generation, so inference computation can grow without adding parameters. On ImageNet at 256x256, LiFT-L/2 achieves an FID 3.34 points lower than our dense DiT-XL/2 baseline while using approximately 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑