发表机构
SK Telecom(SK电信)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
A.X K2 是参数量 6880 亿的 MoE 语言模型,针对智能体应用优化,采用 SGA、GN 等技术,在多基准测试中优于前身,适配成本低,支持长上下文,与开源基准竞争力相当。
AI 中文摘要
我们推出 A.X K2,这是一个参数量达 6880 亿的混合专家(MoE)语言模型,从头开始训练,作为高性能基础模型用于智能体应用。该模型在约 8.5 万亿个 token 上训练,少于其前身 A.X K1,训练数据是规模更小但质量更高的混合数据,其中大幅扩展了智能体和软件工程相关数据;尽管如此,它在所有基准测试中均优于 A.X K1,部分基准测试的提升超过 30 个百分点,体现出 token 效率的大幅提升。为高效支持长上下文,我们推出稀疏门控注意力(SGA),它结合了稀疏注意力与门控注意力,并采用门控归一化(GN)来稳定大规模训练。SGA 通过稀疏索引器预热原生训练至 128K,该预热过程针对索引器自身的稀疏 top-k 选择而非密集注意力分布进行优化,使适配成本显著降低:每个查询仅读取 2048 个位置,但长上下文质量保持不变,且 A.X K2 在 RULER 基准上达到 94.6 分,支持最长 256K 上下文。GN 的离群值抑制则使 4 位 NVFP4 Serving 的精度与 FP8 精度的差距控制在 1 个点以内。一个简单且有效的“思考融合(Think-Fusion)”方案进一步让用户可在单个统一模型内切换思考模式与非思考模式。大量评估表明,A.X K2 与强大的开源权重基准模型相比具有竞争力,在数学和韩语基准上与它们相当或表现更优。
英文摘要
We introduce A.X K2, a 688B-parameter Mixture-of-Experts (MoE) language model trained from scratch as a high-performance foundation for \emph{agentic} applications. Trained on approximately 8.5T tokens---fewer than its predecessor, A.X K1---on a smaller but higher-quality mixture with substantially expanded agentic and software-engineering data, it nonetheless improves over A.X K1 across the board, by over 30 percentage points on some benchmarks, reflecting large gains in token efficiency. To support long contexts efficiently, we introduce Sparse Gated Attention (SGA), which combines sparse attention with gated attention, and adopt Gated Norm (GN) to stabilize large-scale training. SGA is trained natively at 128K through a \emph{sparse} indexer warmup that optimizes the indexer against its own sparse top-$k$ selection rather than the dense attention distribution, making adaptation markedly cheaper: each query reads only 2,048 positions, yet long-context quality is unchanged and A.X K2 scores 94.6 on RULER out to 256K. The outlier suppression of GN in turn keeps 4-bit NVFP4 serving within one point of FP8 accuracy. A simple yet effective Think-Fusion recipe further lets users switch between thinking and non-thinking modes within a single unified model. Extensive evaluations show that A.X K2 performs competitively against strong open-weight baselines, matching or exceeding them on math and Korean-language benchmarks.
Commentshttps://huggingface.co/skt/A.X-K2