Zarya:一种具有灵活训练与双模式推理的混合自回归-掩码扩散语言模型
Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference
AI总结:
Zarya提出混合自回归-掩码扩散语言模型,通过槽式课程训练和双模式推理,兼顾生成质量与并行效率,实现KV缓存重用。
AI中文摘要:
自回归语言模型受限于顺序的、从左到右的生成方式,而掩码扩散模型支持并行解码,但由于无法重用键值缓存而遭受高计算开销,并且由于在难以处理的token组合空间上学习依赖关系而导致生成不连贯。我们引入了Zarya,一个混合语言模型系列,在单一架构内联合优化自回归目标和掩码扩散目标。Zarya将训练数据组织为可变大小的槽,并采用逐步增加槽粒度的课程学习,从而实现从细粒度自回归学习到粗粒度扩散学习的平滑过渡。在推理时,Zarya通过统一接口提供两种不同的解码范式:(i)带首次命中去噪的MDM采样,以及(ii)槽式投机解码,该解码将槽间基于扩散的选择与槽内自回归填充交错进行,实现完全的键值缓存重用。训练和推理机制完全解耦,允许以任何配置训练的模型以任一模式部署。广泛的配置性——包括分组噪声模式(前缀完成、前缀内填充、中间填充)、有序采样调度和噪声级别排列策略——支持灵活的研究探索。我们公开发布了0.6B、1.7B和4B大小的Zarya模型,在标准基准上展示了性能,同时提供了自回归和扩散范式的原则性集成。
英文摘要:
Autoregressive language models (ARMs) are constrained by sequential, left-to-right generation, while masked diffusion models (MDMs) enable parallel decoding but suffer from high computational overhead due to the inability to reuse Key-Value (KV) cache and from incoherent generation arising from learning dependencies over an intractable space of token combinations. We introduce Zarya, a family of hybrid language models that jointly optimizes an autoregressive (AR) objective and a masked-diffusion objective within a single architecture. Zarya structures training data into variable-size slots and employs a curriculum that gradually increases slot granularity, enabling a smooth transition from fine-grained AR learning to coarse-grained diffusion learning. At inference, Zarya provides two distinct decoding paradigms through a unified interface: (i) MDM sampling with first-hitting denoising, and (ii) slotted speculative decoding that interleaves inter-slot diffusion-based selection with intra-slot autoregressive infilling, achieving full KV cache reuse. The training and inference regimes are fully decoupled, allowing a model trained with any configuration to be deployed in either mode. Extensive configurability --- including grouped noise patterns (Prefix Completion, Fill-In-the-Prefix, Fill-In-the-Middle), ordered sampling schedules, and noise-level permutation strategies --- enables flexible research exploration. We release Zarya models publicly in sizes 0.6B, 1.7B, and 4B, demonstrating performance on standard benchmarks while offering a principled integration of autoregressive and diffusion paradigms.