arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10715cs.CL

NCP-ArchPreview 技术报告:通过下一概念预测迈向潜空间语言模型

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

NCP Team, Jiaqi Cao, Chiyu Chen, Shuang Cheng, Xu Cheng, Beiya Dai, Yufan Feng, Kewen Ge, Ruijun Ge, Jiayi Huang, Yang Jiao, Dahua Lin, Zhouhan Lin, Yifan Liu, … 展开作者

NCP Team, Jiaqi Cao, Chiyu Chen, Shuang Cheng, Xu Cheng, Beiya Dai, Yufan Feng, Kewen Ge, Ruijun Ge, Jiayi Huang, Yang Jiao, Dahua Lin, Zhouhan Lin, Yifan Liu, Yuliang Liu, Biqing Qi, Mowen Ruan, Junzhe Shen, Yunchong Song, Hao Sun, Zhongbo Tian, Yixuan Wang, Rubin Wei, Jiaxin Xiong, Kangyu Yang, Qian Yao, Qi Zhang, Bowen Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

提出潜空间语言模型NCP-ArchPreview,通过下一概念预测联合训练,在8.9B参数和5.73T词元上超越OLMo-3-7B,并展示轻量级领域自适应和加速解码优势。

中文摘要 AI 辅助

我们介绍了 NCP-ArchPreview,一种潜空间语言模型,它将自回归预训练推向了标准下一词元预测(NTP)之外。在 NTP 之外,该模型还通过下一概念预测(NCP)学习预测跨越多个词元的离散概念,引入了显式且更具挑战性的概念级目标,同时保留了标准的词元级自回归生成。NCP-ArchPreview 通过直接从其隐藏状态构建乘积量化概念词汇表来构建潜空间,随后通过专门的概念模块学习预测未来概念。这些预测的概念随后被反馈到词元级别以指导后续生成,NTP 和 NCP 以端到端方式联合训练。我们将此架构扩展到 8.9B 参数,并在来自 Dolma-3 数据集的 5.73T 词元上训练,标志着迄今为止最大规模的潜空间语言模型演示。值得注意的是,仅消耗总训练词元的 51.3%,NCP-ArchPreview 就达到了 OLMo-3-7B 的最终预训练损失。在完整预训练之后,它在下游宏平均上比 OLMo-3-7B 高出 2.45 分,包括在 GSM8K 上显著的 5.99 分提升。受控实验分离了来自潜架构和 NCP 目标的性能提升的清晰进展。此外,仅使用标准计算的 85%,NCP-ArchPreview 接近严格参数对齐的 8.9B 基线的训练损失。学习到的潜空间在预训练阶段之后仍然极具价值:仅更新 17M 参数的 VQ 模块就产生了一种新颖的轻量级领域自适应接口,而将概念表示简单注入 DFlash2 起草器可将平均接受长度提高 4.17%,且开销可忽略不计。

英文摘要

We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and more challenging concept-level objective while preserving standard token-level autoregressive generation. NCP-ArchPreview builds a latent space by constructing a product-quantized concept vocabulary directly from its hidden states, and subsequently learns to predict future concepts via a dedicated Concept Module. These predicted concepts are then fed back to the token level to guide subsequent generation, with NTP and NCP trained jointly end-to-end. We scale this architecture to 8.9B parameters and train it on 5.73T tokens from the Dolma-3 dataset, marking the largest demonstration of a latent-space language model to date. Remarkably, by consuming only 51.3% of the total training tokens, NCP-ArchPreview achieves the final pretraining loss of OLMo-3-7B. Following full pretraining, it outperforms OLMo-3-7B by 2.45 points on the downstream macro-average, including a notable 5.99-point gain on GSM8K. Controlled experiments isolate a clear progression of performance gains stemming from both the latent architecture and the NCP objective. Furthermore, utilizing only 85% of the standard computation, NCP-ArchPreview approaches the training loss of a strictly parameter-aligned 8.9B baseline. The learned latent space remains highly valuable after the pretraining stage: updating just the 17M-parameter VQ module yields a novel, lightweight interface for domain adaptation, while a simple injection of concept representations into a DFlash2 drafter improves the mean accepted length by 4.17% with negligible overhead.

发表机构

  • Shanghai AI Lab(上海人工智能实验室)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑