NCP-ArchPreview 技术报告:通过下一概念预测迈向潜空间语言模型
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
浏览论文内容
中文总结 AI 辅助
提出潜空间语言模型NCP-ArchPreview,通过下一概念预测联合训练,在8.9B参数和5.73T词元上超越OLMo-3-7B,并展示轻量级领域自适应和加速解码优势。
中文摘要 AI 辅助
我们介绍了 NCP-ArchPreview,一种潜空间语言模型,它将自回归预训练推向了标准下一词元预测(NTP)之外。在 NTP 之外,该模型还通过下一概念预测(NCP)学习预测跨越多个词元的离散概念,引入了显式且更具挑战性的概念级目标,同时保留了标准的词元级自回归生成。NCP-ArchPreview 通过直接从其隐藏状态构建乘积量化概念词汇表来构建潜空间,随后通过专门的概念模块学习预测未来概念。这些预测的概念随后被反馈到词元级别以指导后续生成,NTP 和 NCP 以端到端方式联合训练。我们将此架构扩展到 8.9B 参数,并在来自 Dolma-3 数据集的 5.73T 词元上训练,标志着迄今为止最大规模的潜空间语言模型演示。值得注意的是,仅消耗总训练词元的 51.3%,NCP-ArchPreview 就达到了 OLMo-3-7B 的最终预训练损失。在完整预训练之后,它在下游宏平均上比 OLMo-3-7B 高出 2.45 分,包括在 GSM8K 上显著的 5.99 分提升。受控实验分离了来自潜架构和 NCP 目标的性能提升的清晰进展。此外,仅使用标准计算的 85%,NCP-ArchPreview 接近严格参数对齐的 8.9B 基线的训练损失。学习到的潜空间在预训练阶段之后仍然极具价值:仅更新 17M 参数的 VQ 模块就产生了一种新颖的轻量级领域自适应接口,而将概念表示简单注入 DFlash2 起草器可将平均接受长度提高 4.17%,且开销可忽略不计。
英文摘要
We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and more challenging concept-level objective while preserving standard token-level autoregressive generation. NCP-ArchPreview builds a latent space by constructing a product-quantized concept vocabulary directly from its hidden states, and subsequently learns to predict future concepts via a dedicated Concept Module. These predicted concepts are then fed back to the token level to guide subsequent generation, with NTP and NCP trained jointly end-to-end. We scale this architecture to 8.9B parameters and train it on 5.73T tokens from the Dolma-3 dataset, marking the largest demonstration of a latent-space language model to date. Remarkably, by consuming only 51.3% of the total training tokens, NCP-ArchPreview achieves the final pretraining loss of OLMo-3-7B. Following full pretraining, it outperforms OLMo-3-7B by 2.45 points on the downstream macro-average, including a notable 5.99-point gain on GSM8K. Controlled experiments isolate a clear progression of performance gains stemming from both the latent architecture and the NCP objective. Furthermore, utilizing only 85% of the standard computation, NCP-ArchPreview approaches the training loss of a strictly parameter-aligned 8.9B baseline. The learned latent space remains highly valuable after the pretraining stage: updating just the 17M-parameter VQ module yields a novel, lightweight interface for domain adaptation, while a simple injection of concept representations into a DFlash2 drafter improves the mean accepted length by 4.17% with negligible overhead.
发表机构
- Shanghai AI Lab(上海人工智能实验室)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。