发表机构
Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU); Fraunhofer IIS(埃尔朗根-纽伦堡大学; 弗劳恩霍夫集成电路研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对生成式复原中CFM方法的时间条件局限,提出判别式流匹配方法,利用判别式流状态表征替代时间条件,在语音增强、图像去噪任务中性能优于基线方法。
AI 中文摘要
现有条件流匹配(CFM)公式通过显式插值坐标(通常被解释为时间)描述传输进度,假设单个全局变量足以表征样本在生成轨迹上的位置。然而在复原任务中,传输进度与样本相关,因为初始分布可能与目标分布呈现不同的统计依赖关系,导致相同插值坐标下的样本在退化程度、到目标分布的距离及复原难度上存在显著差异。本研究探究判别式训练模型学习到的信号表征是否能为基于CFM的复原任务提供有意义的生成传输状态描述。通过系统的隐空间分析,研究表明判别式表征按退化严重程度组织,并在生成过程中遵循指向干净数据流形的一致轨迹。基于上述观察,提出判别式流状态假设,即判别式表征编码控制生成式复原的传输状态。基于该假设,提出判别式流匹配,其将流匹配速度场的条件设置为判别式流状态表征,而非显式时间坐标。在语音增强和图像去噪实验中,结果显示这些表征可表征复原进度、支持自适应推理,且始终优于CFM及扩散相关基线。研究发现表明,判别式表征为显式时间条件提供了有效的状态感知替代方案,并为判别式建模与基于CFM的生成式建模之间的关系提供了新视角。
英文摘要
Existing Conditional Flow Matching (CFM) formulations describe transport progress using an explicit interpolation coordinate, commonly interpreted as time, assuming that a single global variable adequately represents a sample's position along the generative trajectory. In restoration tasks, however, transport progress is sample-dependent because the initial distribution may exhibit varying statistical dependencies with the target distribution. Thus, samples at the same interpolation coordinate can differ substantially in degradation level, distance to the target distribution, and restoration difficulty. We investigate whether signal representations learned by discriminatively trained models provide a meaningful description of generative transport state in CFM-based restoration. Through systematic latent-space analysis, we show that discriminative representations organize according to degradation severity and follow a consistent trajectory toward the clean-data manifold during generation. Motivated by these observations, we introduce the Discriminative Flow-State Hypothesis, which posits that discriminative representations encode a transport state governing generative restoration. Based on this hypothesis, we propose Discriminative Flow Matching, which conditions the Flow-Matching velocity field on Discriminative Flow-State Representations rather than explicit time coordinates. Experiments on speech enhancement and image denoising show that these representations characterize restoration progress, enable adaptive inference, and consistently outperform CFM and diffusion-related baselines. Our findings suggest that discriminative representations provide an effective state-aware alternative to explicit time conditioning and offer a novel perspective on the relationship between discriminative and CFM-based generative modeling.