TuringLLM:面向物理AI高效扩展基础模型
TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI
查看机构详情
- Xpeng Inc(小鹏汽车)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出Turing-20B-A2B混合专家语言模型,采用分位数路由、混合注意力等技术,在紧凑激活参数下实现模型能力、长上下文扩展与推理效率的平衡,性能优于Qwen3-8B Base且接近Qwen3.5-9B Base。
中文摘要 AI 辅助
我们提出Turing-20B-A2B,这是一个拥有200亿参数的混合专家(Mixture-of-Experts)语言模型,每个token仅激活约20亿参数,专为长上下文且对延迟敏感的物理AI应用设计。该模型采用分位数路由(Quantile Routing)的动态top-k配置,实现token自适应的专家分配,同时保持专家利用率均衡并控制平均计算预算。部署阶段,我们进一步对提示词预填充应用容量受限路由,以实现更规则高效的专家执行,而预训练阶段则保留无丢弃(dropless)路由。Turing-20B-A2B还采用混合注意力架构,结合闪电注意力(Lightning Attention)与少量全注意力层,实现高效长上下文建模。该模型通过渐进式三阶段课程进行预训练,并通过持续预训练扩展至原生128K上下文长度,利用YaRN进一步将推理时上下文扩展至512K。尽管激活参数预算紧凑,Turing-20B-A2B在基础模型阶段的整体通用能力超过Qwen3-8B Base,接近Qwen3.5-9B Base,同时保持强劲的长上下文性能和良好的预填充延迟扩展能力。这些结果证明了模型能力、长上下文可扩展性与实际推理效率之间的有效平衡。
英文摘要
We present Turing-20B-A2B, a 20B-parameter Mixture-of-Experts language model that activates approximately 2B parameters per token, designed for long-context and latency-sensitive physical AI applications. The model adopts Quantile Routing in a dynamic top-k configuration, enabling token-adaptive expert allocation while maintaining balanced expert utilization and a controlled average compute budget. During deployment, we further apply capacity-constrained routing to prompt prefill for more regular and efficient expert execution, while retaining dropless routing during pretraining. Turing-20B-A2B also employs a hybrid attention architecture that combines Lightning Attention with a small number of full-attention layers for efficient long-context modeling. The model is pretrained with a progressive three-stage curriculum and extended to a native context length of 128K through continued pretraining, with further inference-time extension to 512K using YaRN. Despite its compact active-parameter budget, Turing-20B-A2B achieves, at the base-model stage, overall general capability exceeding Qwen3-8B Base and approaching Qwen3.5-9B Base, while maintaining strong long-context performance and favorable prefill-latency scaling. These results demonstrate an effective balance among model capability, long-context scalability, and practical inference efficiency.