发表机构
Johns Hopkins University; National University of Singapore(约翰斯·霍普金斯大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究重尾数据匹配问题,提出HTFM框架,将重尾源视为时钟条件高斯源混合,用截断对数签名特征编码时钟。实验显示其在多领域优于高斯流匹配等基线,保留低NFE采样优势,还提供尾部控制接口。
AI 中文摘要
重尾数据出现在许多领域,如不平衡图像数据集、金融回报和极端天气等,其中罕见事件具有不成比例的重要性。标准扩散和流匹配模型通常从高斯噪声或高斯源分布开始,对重尾数据的归纳匹配较差。我们提出了通过随机时钟进行重尾流匹配(HTFM)框架,将重尾源描绘为时钟条件高斯源的混合。给定时钟路径时,源分布和流是高斯的;对时钟求边缘分布得到覆盖高斯、α稳定和学生t族的高斯尺度混合。为使时钟条件向量场实用,我们使用截断对数签名特征编码路径值时钟,使速度场能以可忽略的开销适应已实现的条件空间。实验表明,在二维不平衡α稳定混合、CIFAR10-LT和HRRR天气场中,HTFM在模式覆盖、样本质量和尾部统计恢复方面优于高斯流匹配和有竞争力的重尾基线,同时保留了流匹配的低NFE采样优势。此外,随机时钟公式还提供了一个实用的尾部控制接口。
英文摘要
Heavy-tailed data arise in many domains where rare events carry disproportionate importance, such as imbalanced image datasets, financial returns, and weather extremes. Standard diffusion and flow-matching models typically begin from Gaussian noise or Gaussian source distributions, which yield tractable training targets but provide a poor inductive match for heavy-tailed data. We propose Heavy-Tailed Flow Matching via Random Clocks (HTFM), a framework that portrays heavy-tailed sources as mixtures of clock-conditioned Gaussian sources. Conditioning on a given clock path, the source distribution and flow are Gaussian; marginalizing over the clock gives a Gaussian scale mixture covering Gaussian, $α$-stable, and Student-t families. To make the clock-conditioned vector field practical, we encode the path-valued clock using truncated logsignature features, allowing the velocity field to adapt to the realized conditional space with negligible overhead. Empirically, on 2D imbalanced $α$-stable mixtures, CIFAR10-LT, and HRRR weather fields, HTFM improves mode coverage, sample quality, and tail-statistic recovery over Gaussian flow matching and competitive heavy-tailed baselines, while retaining the low-NFE sampling advantage of flow matching. Moreover, the random-clock formulation further provides a practical tail-control interface: by varying only the clock law or tail parameter, the same architecture can calibrate the ``heaviness'' of generated tails across different distribution families.