arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

上下文流匹配:流模型中的自适应步长选择用于高效视觉生成

Contextual Flow Matching: Adaptive Step Selection in Flow Models for Efficient Visual Generation

Divya Jyoti Bajpai, Arun Verma, Manjesh Kumar Hanawal

arXiv 2610.03202首次发表:更新:

发表机构

IIT Bombay; SMART Centre(印度理工学院孟买分校; SMART中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出COFLOW,一种推理时自适应步长选择方法,基于提示特征平衡效率与保真度,无需重训模型,在图像和视频生成中实现超2.5倍加速并保持质量。

AI 中文摘要

流匹配通过连续时间动力学实现高质量视觉生成,但推理过程因多次顺序函数评估而成本高昂。现有加速方法减少了函数评估次数,但往往引入额外训练开销、降低质量或未能考虑输入依赖性变化。我们提出COFLOW,一种推理时方法,基于提示特征自适应地为每次生成选择步数。我们的上下文感知COFLOW通过无监督奖励在线训练,该奖励平衡推理效率与生成保真度。我们的方法即插即用,无需重新训练底层生成模型。它适用于图像和视频生成,在保持感知和语义质量的同时实现超过2.5倍的加速。我们进一步提供理论分析,在标准正则性条件下建立了O(1/K)的前向欧拉离散化误差界。

英文摘要

Flow Matching enables high-quality visual generation via continuous-time dynamics, but inference remains costly due to multiple sequential function evaluations. Existing acceleration methods reduce the number of function evaluations but often introduce additional training overhead, degrade quality, or fail to account for input-dependent variability. We propose COFLOW, an inference-time method that adaptively selects the step counts each generation based on the prompt features. Our context-aware COFLOW is trained online with an unsupervised reward that balances inference efficiency and generation fidelity. Our method is plug-and-play, requiring no retraining of the underlying generative model. It generalizes to image and video generation, achieving over 2.5x speedup while preserving perceptual and semantic quality. We further provide a theoretical analysis establishing an O(1/K) forward-Euler discretization error bound under standard regularity conditions.

CommentsAccepted in NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑