arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过非绑定自条件化改进少步长语言流模型

Improving Few-Step Language Flows with Untied Self-Conditioning

Bocheng Li, Linli Xu

arXiv 2608.22244首次发表:更新:

发表机构

University of Science and Technology of China; State Key Laboratory of Cognitive Intelligence(中国科学技术大学; 认知智能国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对少步长流匹配语言模型的训练-推理不匹配问题,提出非绑定自条件化采样器,无需重训即可在多数据集上显著降低生成困惑度、提升生成偏好度。

AI 中文摘要

流匹配语言模型可并行优化所有 token 位置,并能在采样步数与延迟间做权衡,但采样步数较少时生成质量仍会急剧下降。我们将该下降的来源追溯至先前预测自条件化中存在的训练-推理不匹配:训练时,自条件化输入由当前含噪状态计算而来,无中间求解器步;采样时,求解器会将前序预测折叠进隐变量,随后该预测才会作为显式自条件化输入出现。训练时不存在的这种耦合会产生冗余,且冗余随步长增大而增长。我们表明该不匹配会同时降低自条件化输入与求解器更新效果,并从模型自身结构中为二者各推导了一种修正方案:从冻结的投影权重中识别出自条件化输入与隐变量冗余的方向并对其进行抑制;从求解器的积分结构中推导得出需要步长平均预测,并从预测历史中对其进行近似,尺度由离线轨迹统计值设定。得到的采样器名为非绑定自条件化,无需重新训练,每步仅需一次评估。在 LangFlow 上使用 8 个采样步时,它将 OpenWebText 生成困惑度从 531 降至 62(降幅达 8.6 倍);在适配后的 Arena-Hard-Auto v2 协议下,其输出在 96% 的成对比较中更受偏好;在 ELF-B 上,它将生成困惑度从 71 降至 43。改进效果在 8 至 256 个采样步范围内均保持一致。

英文摘要

Flow-matching language models refine all token positions in parallel and can trade sampling steps for latency, yet generation quality still degrades sharply with few sampling steps. We trace a source of this degradation to a train--inference mismatch in previous-prediction self-conditioning: during training, the self-conditioning input is computed from the current noisy state with no intervening solver step; during sampling, the solver folds the previous prediction into the latent before that same prediction reappears as the explicit self-conditioning input. This coupling, absent during training, creates redundancy that grows with step width. We show that the mismatch degrades both the self-conditioning input and the solver update, and derive a correction for each from the model's own structure. From the frozen projection weights we identify directions along which the self-conditioning input is redundant with the latent and dampen them; from the solver's integration structure we derive that a step-average prediction is needed and approximate it from prediction history, with scale set by offline trajectory statistics. The resulting sampler, Untied Self-Conditioning, requires no retraining and uses one evaluation per step. At 8 sampling steps on LangFlow, it reduces OpenWebText generative perplexity from $531$ to~$62$ ($8.6\times$); under an adapted Arena-Hard-Auto~v2 protocol, its outputs are preferred in $96\%$ of pairwise comparisons. On ELF-B it reduces generative perplexity from $71$ to~$43$. Improvements hold from 8 to 256 sampling steps.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑