无需训练的流匹配图像生成器的隐状态优化
Training-Free Hidden-State Refinement for Flow-Matching Image Generators
- Huazhong University of Science and Technology(华中科技大学)
- Tongji University(同济大学)
- King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
- University of California, Merced(加州大学默塞德分校)
- University of Amsterdam(阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究提出无需训练的循环框架优化流匹配图像生成器,在不改动权重与采样器的情况下提升质量指标,Loop Guidance在Scale-RAE DiT2.4B上显著提升GenEval与DPG-Bench指标。
中文摘要 AI 辅助
我们旨在通过在去噪器内部添加推理计算来改进冻结的流匹配图像生成器,而无需修改模型权重或外部采样器。现有生成器通常通过增加采样步数来花费额外的测试时间计算,这会反复评估整个去噪器,且质量提升与采样器成本挂钩。一个关键挑战是如何在冻结的Transformer去噪器内部使用额外计算:该方法必须决定哪些token、层和采样时间应接收重复更新,同时保留原始生成流程。我们提出了一种无需训练的循环框架,该框架在每次去噪调用中重复应用选定的Transformer层。Dense Token Loop(密集token循环)和Sparse Token Loop(稀疏token循环)在token范围上有所不同;Sampling-Progress Gating(采样进度门控)和循环层范围指定循环何时及何处激活;循环次数和强度控制重复更新;Loop Guidance(循环引导)结合普通和循环的向量场预测。在两个Scale-RAE模型规模上,循环变体提升了主要和辅助质量指标,且具有竞争力的质量-效率权衡;Loop Guidance进一步提升了所有三个测试模型的主要指标,在Scale-RAE DiT2.4B上,其将GenEval从0.4471提升至0.5691,将DPG-Bench从0.7656提升至0.8053。代码将被发布。
英文摘要
We aim to improve frozen flow-matching image generators by adding inference computation inside the denoiser, without changing model weights or the outer sampler. Existing generators usually spend extra test-time computation by increasing the number of sampling steps, which repeatedly evaluates the entire denoiser and couples quality gains to sampler cost. A key challenge is how to use extra computation inside a frozen transformer denoiser: the method must decide which tokens, layers, and sampling times receive repeated updates while preserving the original generation pipeline. We introduce a training-free looping framework that repeatedly applies selected transformer layers inside each denoising call. Dense and Sparse Token Loop vary the token scope; Sampling-Progress Gating and the loop layer range specify when and where looping is active; loop count and strength control the repeated updates; and Loop Guidance combines ordinary and looped vector-field predictions. Across two Scale-RAE model scales, loop variants improve primary and auxiliary quality metrics with competitive quality--efficiency trade-offs. Loop Guidance further improves both primary metrics across all three tested models; on Scale-RAE DiT2.4B, it raises GenEval from 0.4471 to 0.5691 and DPG-Bench from 0.7656 to 0.8053. Code will be released.