ReAL:通过共享前瞻的分段推进加速流匹配
ReAL: Accelerating Flow Matching through Segment Advancement with Shared Lookahead
AI总结:
ReAL是一种无需训练的流匹配加速采样器,通过共享前瞻选择推进跨度,仅需一次评估即可加速图像和视频生成,在FLUX.1-dev上实现4.91倍加速并保持97%的奖励分数。
AI中文摘要:
流匹配模型能够生成高质量的图像和视频,但重复的神经网络评估使得采样成本高昂。跳过评估可通过在更长的跨度上扩展可用的速度估计来降低这一成本。然而,局部的速度一致性本身并不能确定合适的跨度,且检查每个候选端点会带来昂贵的模型调用。我们引入了ReAL,一种无需训练的采样器,它利用一个共享的前瞻来选择推进的距离。我们的关键洞察在于,未校正与前瞻校正的端点提议之间的差异可以直接从观测到的速度失配和前瞻之外的候选跨度计算得出。这一关系提供了一个依赖于跨度的选择标准,而无需额外的端点评估。同一个前瞻选择最长的可通过候选跨度,校正接受的更新,并将其速度作为下一步的起始估计。初始化后,每次常规迭代仅需一次新的评估。ReAL使用预训练的速度输出和原始噪声调度,无需额外训练或访问内部特征。实验覆盖了四个图像生成骨干网络、视频生成和图像编辑。ReAL在FLUX.1-dev上实现了4.91倍的实测加速,同时保留了密集平均ImageReward的97.0%。在HunyuanVideo上,它实现了5.49倍的加速,同时保持了与密集采样相近的VBench分数。
英文摘要:
Flow-matching models generate high-quality images and videos, but repeated neural network evaluations make sampling expensive. Skipping evaluations reduces this cost by extending an available velocity estimate over a longer span. However, local velocity agreement alone does not determine a suitable span, and checking each candidate endpoint adds costly model calls. We introduce ReAL, a training-free sampler that selects how far to advance using one shared lookahead. Our key insight is that the discrepancy between uncorrected and lookahead-corrected endpoint proposals can be computed directly from the observed velocity mismatch and the candidate span beyond the lookahead. This relation provides a span-dependent selection criterion without additional endpoint evaluations. The same lookahead selects the longest passing candidate span, corrects the accepted update, and supplies its velocity as the next starting estimate. After initialization, each regular iteration requires only one fresh evaluation. ReAL uses the pretrained velocity output and original noise schedule, with no additional training or access to internal features. Experiments cover four image-generation backbones, video generation, and image editing. ReAL achieves 4.91x measured speedup on FLUX.1-dev while retaining 97.0% of dense mean ImageReward. On HunyuanVideo, it achieves a 5.49x speedup while maintaining a VBench score close to that of dense sampling.