arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23551cs.CLcs.AIcs.LGstat.ML

ConvergeFlow:可收敛至词元嵌入的语言流模型

ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings

  • Chinese University of Hong Kong(香港中文大学)
  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

Na Li, Yuchen Jiao, Changxiao Cai, Gen Li

AI总结:

本研究提出ConvergeFlow,一种嵌入空间的基于流的语言模型,通过约束数据预测器至词元嵌入凸包实现可收敛性,无需交叉熵解码器,在OpenWebText上取得与现有模型相当的性能,展现了基于流范式的语言建模潜力。

AI中文摘要:

连续扩散模型和基于流的语言模型(LMs)的近期进展已取得与离散语言模型相当的性能。然而,现有连续框架仍依赖交叉熵(CE)监督的解码器,因为流轨迹无法保证终止于有效词元嵌入。受此局限启发,我们提出ConvergeFlow,一种嵌入空间的基于流的语言模型,它将数据预测器约束在词元嵌入的凸包内,仅用流匹配产生的均方误差目标进行训练。在适当的正则条件下,我们证明即便数据预测器存在误差,所得流仍会收敛至有效词元嵌入,从而无需CE监督解码器即可直接进行词元预测。我们还开发了三种采样机制,用于控制生成困惑度与熵之间的权衡。在OpenWebText上的实验表明,ConvergeFlow取得了与现有连续和离散扩散语言模型相当的性能。这些发现证明了基于流的范式在语言建模中的潜力,我们的代码可在该网址获取。

英文摘要:

Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive with discrete LMs. However, existing continuous frameworks still rely on decoders supervised with cross entropy (CE) because the flow trajectories are not guaranteed to terminate at valid token embeddings. Motivated by this limitation, we introduce \textbf{ConvergeFlow}, an embedding-space flow-based LM, which constrains the data predictor to the convex hull of token embeddings and trains it solely with the mean squared error objective induced by flow matching. Under suitable regularity conditions, we prove that the resulting flow converges to valid token embeddings despite errors in the data predictor, enabling direct token prediction without a CE-supervised decoder. We further develop three sampling mechanisms for controlling the trade-off between the generative perplexity and entropy. Experiments on OpenWebText demonstrate that ConvergeFlow achieves performance competitive with existing continuous and discrete diffusion LMs. These findings demonstrate the potential of the flow-based paradigm for language modeling. Our code is available at https://github.com/Na-Li66/ConvergeFlow.

↑