发表机构
KAIST; Sony Group Corporation; The University of Tokyo(韩国科学技术院; 索尼集团公司; 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文发现均匀离散扩散模型中的显式时间条件化在实践中常非必要,尽管理论上最优预测依赖时间,但在有限数据下经验最优预测对时间不敏感,时间无关预测器可媲美甚至超越时间条件模型。
AI 中文摘要
均匀离散扩散模型(UDMs)通常使用显式的时间条件化,但我们发现这在实践中往往是不必要的。在本文中,我们首先表明,总体最优的UDM预测器通常依赖于时间:时间控制着模型应多大程度信任观察到的上下文。然后我们表明,在与语言相关的有限数据设置中,这种依赖性可能变得可以忽略不计。当一条被破坏的训练序列仍比竞争训练序列更接近其原始干净序列时,经验最优预测器在扩散轨迹的大部分区域对时间几乎不敏感,而这种保证在接近高噪声端点时会减弱。在实证中,经过训练的语言UDM在轨迹的大部分区域表现出有限的时间敏感性,而时间无关的预测器在跨数据集和训练目标上与时间条件模型保持竞争力,并且常常优于后者。这些结果对UDM中显式时间条件化的使用提出了挑战:尽管总体最优依赖于时间,但在实践中显式地以时间为条件可能往往是不必要的。
英文摘要
Uniform discrete diffusion models (UDMs) commonly use explicit time conditioning, but we find that it can often be unnecessary in practice. In this paper, we first show that the population-optimal UDM predictor generally depends on time: time controls how much the model should trust the observed context. We then show that this dependence can become negligible in finite-data settings relevant to language. When a corrupted training sequence remains much closer to its original clean sequence than to competing training sequences, the empirical-optimal predictor is nearly insensitive to time over most of the diffusion trajectory, where the guarantee weakens toward the high-noise endpoint. Empirically, trained language UDMs exhibit limited time sensitivity over most of the trajectory, while time-agnostic predictors remain competitive with, and often outperform, time-conditioned models across datasets and training objectives. These results challenge the use of explicit time conditioning in UDMs: although the population optimum depends on time, explicitly conditioning on it may often be unnecessary in practice.
CommentsPreprint