发表机构
Tencent Hunyuan Team(腾讯混元团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文介绍了 Hunyuan-A13B,一个基于混合专家架构的开源大语言模型,总参数 800 亿、激活 130 亿,通过增强 STEM 数据预训练和双模式思维链框架,在数学、科学、编程等任务上接近更大模型性能,并支持高吞吐推理。
AI 中文摘要
我们推出了 Hunyuan-A13B,一个基于混合专家(Mixture-of-Experts)架构的开源大语言模型。该模型总参数量为 800 亿,但在推理时仅激活 130 亿参数,从而在模型能力、计算效率和部署成本之间取得了平衡。模型在一个经过严格过滤的 20T token 语料库上进行了预训练,并增强了 STEM 数据的策展,以提高事实可靠性和推理能力。高质量的有监督微调和大规模强化学习进一步提升了其整体性能。Hunyuan-A13B 还引入了双模式思维链(Chain-of-Thought)框架,根据任务复杂度调整推理深度:对于常规查询采用快速思考,对于复杂的多步骤问题采用慢速思考。评估显示,该模型在数学、科学、编程、通用语言理解和智能体任务上表现出具有竞争力的性能,通常接近更大规模模型的表现。其高推理吞吐量使其适用于对延迟敏感的应用。我们发布了 Hunyuan-A13B,以支持开放研究和实际的 LLM 部署。
英文摘要
We present Hunyuan-A13B, an open-source large language model based on a Mixture-of-Experts architecture. It contains 80 billion total parameters but activates only 13 billion during inference, balancing model capability, computational efficiency, and deployment cost. The model is pretrained on a rigorously filtered 20T-token corpus with enhanced STEM data curation, improving factual reliability and reasoning ability. High-quality supervised fine-tuning and large-scale reinforcement learning further enhance its overall performance. Hunyuan-A13B also introduces a dual-mode Chain-of-Thought framework that adapts reasoning depth to task complexity: fast thinking for routine queries and slow thinking for complex, multi-step problems. Evaluations show competitive performance across mathematics, science, programming, general language understanding, and agent tasks, often approaching that of much larger models. Its high inference throughput makes it suitable for latency-sensitive applications. We release Hunyuan-A13B to support open research and practical LLM deployment.