发表机构
KAIST; University of Amsterdam; Carnegie Mellon University(韩国科学技术院; 阿姆斯特丹大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出不动点流框架,将自条件机制解释为不动点迭代,并设计二维流模型同时压缩迭代与流过程,实现少步生成,在OpenWebText上超越现有模型。
AI 中文摘要
自条件是一种增强连续流式语言模型的核心技术,模型通过条件化自身的去噪估计来学习去噪生成的文本。尽管经验上成功,但其性能提升的原因尚不明确。此外,基于流映射的少步生成器越来越受关注,但如何利用自条件尚不清楚。本文证明,具有自条件的流语言模型求解了一个不动点迭代,该迭代提升了学习去噪器的性能。我们利用这一观点提出了不动点流,一种二维的自条件流,其中第一维表示流过程,第二维表示不动点迭代。我们证明不动点流定义了有效的流映射,并表明可以通过压缩不动点迭代和流过程从自条件流模型中蒸馏得到,前者通过不动点蒸馏,后者通过流映射蒸馏。由此得到的流映射语言模型FMLM$^\star$在OpenWebText上的单步和少步生成中优于最先进的自条件模型和少步模型。代码可在该https URL获取。
英文摘要
Self-conditioning is a core technique that enhances continuous flow-based language models, where the model learns to denoise generated text by conditioning on its own denoising estimate. While empirically successful, its performance improvements are poorly understood. Moreover, there is growing interest in the use of few-step generators based on flow maps, for which how to leverage self-conditioning is unclear. Here, we show that flow language models with self-conditioning perform a fixed-point iteration that improves generation through iterative refinement. We use this viewpoint to formulate fixed-point flows, a two-dimensional class of self-conditioned flows, where the first dimension represents the flow process and the second represents the fixed-point iteration. We show that fixed-point flows define valid flow maps, and show that they can be distilled from self-conditioned flow models by compressing both fixed-point iterations and the flow process, the former with fixed-point distillation and the latter with flow map distillation. Our resulting flow map language model, FMLM$^\star$, outperforms state-of-the-art self-conditioned models and few-step models in one- and few-step generation on OpenWebText. Code is available at https://github.com/Ugness/self-conditioned-fmlm.