SEA-LION-v4.8:技术报告
SEA-LION-v4.8: A Technical Report
浏览论文内容
中文总结 AI 辅助
本报告介绍基于NVIDIA Nemotron 3的SEA-LION-v4.8模型家族,通过多语言数据适配与后训练,在SEA-HELM上显著提升东南亚语言任务性能。
中文摘要 AI 辅助
我们推出了Nemotron-SEA-LION-v4.8,这是一个基于NVIDIA Nemotron 3构建的东南亚语言统一网络(SEA-LION)模型家族。该家族包括30B-A3B和120B-A12B两种模型,同时提供持续预训练的基础检查点和后训练变体。我们使用东南亚语言、推理、代码和多语言并行数据集对模型进行适配,随后通过监督微调和在线同策略蒸馏进行后训练。在SEA-HELM基准上,30B-A3B模型将整体东南亚语言得分从46.06提升至51.57,而120B-A12B模型则从49.30提升至63.44。在七种东南亚语言的指令遵循、自然语言推理和自然语言理解方面观察到了最显著的提升。
英文摘要
We introduce Nemotron-SEA-LION-v4.8, a family of Southeast Asian Languages In One Network (SEA-LION) models built upon NVIDIA Nemotron 3. The family includes 30B-A3B and 120B-A12B models, with both continued-pretrained base checkpoints and post-trained variants. We adapt the models using Southeast Asian, reasoning, code, and multilingual parallel data, followed by post-training with supervised fine-tuning and online on-policy distillation. On SEA-HELM, the 30B-A3B model improves the overall SEA score from 46.06 to 51.57, while the 120B-A12B model improves from 49.30 to 63.44. Across seven Southeast Asian languages, we observe broad capability gains with the 120B-A12B model showing broader and more consistent improvements across tasks.
发表机构
- AI Singapore(新加坡全国人工智能计划)
机构由 AI 辅助整理,请以论文原文为准。