AI 中文总结
DeltaFlow是噪声自适应双向门控Delta网络,通过两种变体及噪声自适应内存控制、计划时间状态一致性技术,在OpenWebText基准上降低困惑度、提升吞吐量,可作为密集注意力的高效替代方案。
AI 中文摘要
嵌入式语言流(ELF)主要依赖全非因果注意力进行迭代去噪,在每个采样步骤都会反复产生二次序列混合成本。门控Delta网络(GDN)提供了一种高效的循环替代方案,但其标准因果公式无法直接捕获ELF所需的双向上下文。我们引入DeltaFlow,一种用于连续语言去噪的噪声自适应双向GDN主干。我们研究了两种变体:DeltaFlow-A,其在各层之间交替扫描方向;DeltaFlow-P,其在每层内执行并行前向和后向扫描。我们进一步引入噪声自适应内存控制和计划时间状态一致性(TSC),以在相邻噪声水平之间稳定隐藏表示。在OpenWebText上,使用32步随机微分方程采样器,DeltaFlow-P将生成的困惑度从全注意力ELF基线的24.218降至21.228,同时保持相当的一元语法熵,训练令牌暴露量为36B,而基线为45B。在仅去噪器基准测试中,当序列长度为16k时,DeltaFlow-P比全注意力基线实现了2.72倍的吞吐量加速。这些结果表明,DeltaFlow是密集注意力用于高效连续语言去噪的有前途替代方案。
英文摘要
Embedded Language Flows (ELF) rely primarily on full non-causal attention for iterative denoising, repeatedly incurring quadratic sequence-mixing cost at each sampling step. Gated Delta Networks (GDNs) provide an efficient recurrent alternative, but their standard causal formulation cannot directly capture the bidirectional context required by ELF. We introduce DeltaFlow, a noise-adaptive bidirectional GDN backbone for continuous language denoising. We study two variants: DeltaFlow-A, which alternates scan directions across layers, and DeltaFlow-P, which performs parallel forward and backward scans within each layer. We further introduce noise-adaptive memory control and scheduled Temporal State Consistency (TSC) to stabilize hidden representations across nearby noise levels. On OpenWebText, using a 32-step stochastic differential equation sampler, DeltaFlow-P reduces generated perplexity from 24.218 for the full-attention ELF baseline to 21.228 while maintaining comparable unigram entropy, with 36B training-token exposure compared with 45B for the baseline. In a denoiser-only benchmark, DeltaFlow-P achieves a 2.72x throughput speedup over the full-attention baseline at a sequence length of 16k. These results show that DeltaFlow is a promising alternative to dense attention for efficient continuous language denoising.