arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30567cs.AI

TuringLLM:面向物理AI高效扩展基础模型

TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI

发表机构小鹏汽车
查看机构详情
  • Xpeng Inc(小鹏汽车)

机构由 AI 辅助整理,请以论文原文为准。

Yuheng Zhang, Yizhao Wang, Da Zhu, Hua Zhou, Yue He, Jiahui Hu, Shaman Tang, Hanlin Chen, Yuhua Wei, Anhua Liu, Shuang Su, Rui Xin, MingYuan Wang, MingHao Li, H… 展开作者

Yuheng Zhang, Yizhao Wang, Da Zhu, Hua Zhou, Yue He, Jiahui Hu, Shaman Tang, Hanlin Chen, Yuhua Wei, Anhua Liu, Shuang Su, Rui Xin, MingYuan Wang, MingHao Li, HaoJie Yang, Siqi Liu, Jianlei Zheng, WeiChao Huang, Qiman Wu, Hang Zhang, HongGou Yang, Xianming Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Turing-20B-A2B混合专家语言模型,采用分位数路由、混合注意力等技术,在紧凑激活参数下实现模型能力、长上下文扩展与推理效率的平衡,性能优于Qwen3-8B Base且接近Qwen3.5-9B Base。

中文摘要 AI 辅助

我们提出Turing-20B-A2B,这是一个拥有200亿参数的混合专家(Mixture-of-Experts)语言模型,每个token仅激活约20亿参数,专为长上下文且对延迟敏感的物理AI应用设计。该模型采用分位数路由(Quantile Routing)的动态top-k配置,实现token自适应的专家分配,同时保持专家利用率均衡并控制平均计算预算。部署阶段,我们进一步对提示词预填充应用容量受限路由,以实现更规则高效的专家执行,而预训练阶段则保留无丢弃(dropless)路由。Turing-20B-A2B还采用混合注意力架构,结合闪电注意力(Lightning Attention)与少量全注意力层,实现高效长上下文建模。该模型通过渐进式三阶段课程进行预训练,并通过持续预训练扩展至原生128K上下文长度,利用YaRN进一步将推理时上下文扩展至512K。尽管激活参数预算紧凑,Turing-20B-A2B在基础模型阶段的整体通用能力超过Qwen3-8B Base,接近Qwen3.5-9B Base,同时保持强劲的长上下文性能和良好的预填充延迟扩展能力。这些结果证明了模型能力、长上下文可扩展性与实际推理效率之间的有效平衡。

英文摘要

We present Turing-20B-A2B, a 20B-parameter Mixture-of-Experts language model that activates approximately 2B parameters per token, designed for long-context and latency-sensitive physical AI applications. The model adopts Quantile Routing in a dynamic top-k configuration, enabling token-adaptive expert allocation while maintaining balanced expert utilization and a controlled average compute budget. During deployment, we further apply capacity-constrained routing to prompt prefill for more regular and efficient expert execution, while retaining dropless routing during pretraining. Turing-20B-A2B also employs a hybrid attention architecture that combines Lightning Attention with a small number of full-attention layers for efficient long-context modeling. The model is pretrained with a progressive three-stage curriculum and extended to a native context length of 128K through continued pretraining, with further inference-time extension to 512K using YaRN. Despite its compact active-parameter budget, Turing-20B-A2B achieves, at the base-model stage, overall general capability exceeding Qwen3-8B Base and approaching Qwen3.5-9B Base, while maintaining strong long-context performance and favorable prefill-latency scaling. These results demonstrate an effective balance among model capability, long-context scalability, and practical inference efficiency.

补充信息

↑