上下文加权离散流匹配
Context-weighted Discrete Flow Matching
浏览论文内容
中文总结 AI 辅助
研究离散生成建模,提出对连续时间马尔可夫链的简单修改,通过上下文加权采样器和缩放交叉熵损失函数,纳入局部上下文信息,提高生成质量与训练效率,在质量上与强基线匹配且能任意顺序生成。
中文摘要 AI 辅助
离散流匹配为离散结构上的生成建模提供了灵活框架。标准分解训练目标使模型面对不同难度目标,混合了条件良好、可预测的令牌与模糊、高熵的令牌。实验表明每个令牌值的不确定性与邻域可用上下文密度密切相关。基于此,对基础连续时间马尔可夫链进行简单修改以纳入局部上下文信息。上下文加权采样器提高了生成质量且计算开销可忽略不计,缩放交叉熵损失函数重新加权训练信号,在OpenWebText上使生成困惑度降低达63%。该方法在质量上与强大的半自回归块扩散基线匹配,且能以任意顺序生成。结果凸显了局部上下文在离散生成建模中的重要作用,简单的上下文感知修改可显著提高采样和训练效率。
英文摘要
Discrete flow matching provides a flexible framework for generative modeling on discrete structures. However, the standard factorized training objective exposes the model to targets of varying difficulty, mixing well-conditioned, predictable tokens with ambiguous, high-entropy ones. We empirically demonstrate that the uncertainty over the value of each token is closely related to the density of available context in its neighborhood. Motivated by this observation, we propose a simple modification to the underlying continuous-time Markov chain (CTMC) that incorporates local context information. Our context-weighted sampler improves generation quality with negligible computational overhead, while our scaled cross-entropy loss function reweights the training signal from different tokens and reduces generative perplexity by up to 63% on OpenWebText. Moreover, our approach matches a strong semi-autoregressive block diffusion baseline in quality while retaining the ability to perform generation in any order. These results highlight the role of local context as an important factor in discrete generative modeling and show that simple context-aware modifications can significantly improve both sampling and training efficiency.
发表机构
- University of Amsterdam(阿姆斯特丹大学)
- Meta FAIR
机构由 AI 辅助整理,请以论文原文为准。