AI 中文总结
提出用于流匹配的单侧分位数耦合(QC-FM),无需小批量传输,在 CIFAR-10 等四组数据集上降低 FID 最多 12.9%,性能优于基线及 OT-CFM。
AI 中文摘要
流匹配通过回归简单源分布与目标数据分布间概率路径的速度场来训练连续时间生成模型。配对源样本与目标样本的耦合对优化和样本质量影响显著,但结构化耦合通常依赖小批量传输或分配过程,其成本随批量大小至少呈二次增长。我们提出轻量型单侧耦合——分位数耦合流匹配(QC-FM):无需匹配两个预采样批次,仅采样数据批次并直接构建每个配对源;沿少量随机正交方向投影的数据秩被映射到高斯分位数,潜在代码则通过条件高斯采样在正交补空间中完成。该构造每个切片为一维,因此耦合无需成对成本矩阵,也无需求解分配问题。我们证明,对于每个抽取的帧,该耦合消除了每个选定切片上的不可约回归方差,使理想流在该切片上完全笔直,同时保持采样先验不变:生成仍从标准高斯分布开始,训练源仅通过切片代码的 copula(连接函数)偏离标准高斯分布,我们对其传输成本进行了界定。训练时,我们将 QC 应用于锚点子集,并使用精确高斯样本填充剩余源槽,在保留 QC 偏差的同时保留了基线耦合的显式信号。在 CIFAR-10、CelebA、FFHQ 和 ImageNet-64 数据集上,QC-FM 在匹配的训练预算下优于基线,将 FID(Fréchet inception 距离)降低最多 12.9%,且在所有四个数据集上均优于 OT-CFM(最优传输耦合流匹配)。这些结果表明,保留投影秩结构是一种简单且可扩展的方法,可在不求解小批量传输问题的情况下,向流匹配耦合中注入有用的几何偏差。
英文摘要
Flow Matching trains continuous-time generative models by regressing the velocity field of a probability path between a simple source distribution and a target data distribution. The coupling that pairs source and target samples strongly affects optimization and sample quality, but structured couplings typically rely on mini-batch transport or assignment procedures whose cost grows at least quadratically in batch size. We propose Quantile Coupling Flow Matching (QC-FM), a lightweight one-sided coupling: rather than matching two pre-sampled batches, it samples only the data batch and constructs each paired source directly. Data ranks projected along a small number of random orthogonal directions are mapped to Gaussian quantiles, and the latent code is completed in the orthogonal complement by conditional Gaussian sampling. The construction is one-dimensional per slice, so the coupling requires no pairwise cost matrix and no assignment to solve. We show that, for each drawn frame, this coupling eliminates the irreducible regression variance along every selected slice and makes the ideal flow exactly straight there, while leaving the sampling prior unchanged: generation still starts from the standard Gaussian, and the training source deviates from it only through the copula of the slice codes, whose transport cost we bound. For training, we apply QC to an anchor subset and complete the remaining source slots with exact Gaussian samples, retaining the QC bias while preserving an explicit signal from the Baseline coupling. Across CIFAR-10, CelebA, FFHQ, and ImageNet-64, QC-FM improves over the Baseline under matched training budgets, reducing FID by up to 12.9%, and outperforms OT-CFM on all four datasets. These results suggest that preserving projected rank structure is a simple and scalable way to inject useful geometric bias into FM couplings without solving a mini-batch transport problem.