arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Jais 2:以阿拉伯语为核心的开源大语言模型家族

Jais 2: A Family of Arabic-Centric Open Large Language Models

Mohamed Anwar, Abed Alhakim Freihat, George Ibrahim, Mostafa Awad, Abdelrahman Sadallah, Gurpreet Gosal, Gokulakrishnan Ramakrishnan, Sarath Chandran, Biswajit Mishra, Rituraj Joshi, Ahmed Frikha, Etienne Goffinet, Abhishek Maiti, Ali El Filali, Sarah AlBarri, Samujjwal Ghosh, Rahul Pal, Parvez Mullah, Awantika Shukla, Sajid siddiki, Samta Kamboj, Onkar Pandit, Sunil Kumar Sahu, AbdelRahman Elbadawy, Amr Mohamed, Ahmad Chamma, Evan Dufraisse, Abdelaziz Bounhar, Dani Bouch, Hadi Abdine, Guokan Shang, Fajri Koto, Yuxia Wang, Zhuohan Xie, Ali Mekky, Rania Elbadry, Sarfraz Ahmad, Momina Ahsan, Omar El Herraoui, Daniil Orel, Hasan Iqbal, Kareem Elzeky, Mervat Abassy, Kareem Elozeiri, Saadeldine Eletter, Farah Atif, Nurdaulet Mukhituly, Haonan Li, Xudong Han, Aaryamonvikram Singh, Zainul Abedien Ahmed Quraishi, Neha Sengupta, Larry Murray, Avraham Sheinin, Joel Hestness, Natalia Vassilieva, Hector Xuguang Ren, Zhengzhong Liu, Michalis Vazirgiannis, Preslav Nakov

arXiv 2608.13580首次发表:更新:

发表机构

Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学基础模型研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Jais 2是MBZUAI等联合开发的以阿拉伯语为核心的开源大语言模型家族,含70B、8B参数变体,在阿拉伯语及文化相关基准上表现领先,已开源并部署为多平台聊天应用,可支撑相关研究开发。

AI 中文摘要

Jais 2是由MBZUAI、Cerebras和Inception联合开发的以阿拉伯语为核心的大语言模型家族,旨在推进以阿拉伯语为核心的语言建模,在本报告评估的阿拉伯语基准及文化相关基准上表现强劲。据我们所知,该家族包含从头训练的最大规模开源阿拉伯语大语言模型,参数达700亿,以及在评估的开源模型中具有竞争力的80亿参数变体。定制的以阿拉伯语为核心的词汇表可实现高效的训练与推理,此外,优化的架构和训练方案可实现极高的计算效率。与同类模型相比,Jais 2的token预算大幅减少,在本报告考虑的基准上实现了强劲的阿拉伯语性能,同时也具有有竞争力的英语结果。这些模型在评估的开源模型中,在OALL2和AraGen上取得了领先结果,在诗歌、宗教、烹饪、解梦等多个与文化相关的阿拉伯语基准上表现出色,在翻译、摘要等通用任务中也表现优异。我们在HuggingFace上以商业许可发布这些模型,Jais 2 70B还作为聊天应用在网页、iOS和Android上发布,它运行在Cerebras硬件上,每秒可生成多达2000个token,在我们的部署环境中支持高吞吐量的以阿拉伯语为核心的聊天服务。通过将规模、语言多样性、文化保真度、开源性和速度相结合,Jais 2提供了一个开源权重的基础,旨在支持以阿拉伯语为核心的大语言模型的进一步研究与开发。

英文摘要

Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report. The family includes, to our knowledge, the largest open Arabic-centric LLM trained from scratch at 70B parameters, and a competitive 8B-parameter variant among the evaluated open models. A custom Arabic-centric vocabulary enables efficient training and inference. In addition, an optimized architecture and training recipe yield highly compute-efficient training. With a substantially smaller token budget than comparable models, Jais 2 achieves strong Arabic performance on the benchmarks considered in this report and competitive English results. The models obtain leading results among the evaluated open models on OALL2 and AraGen. They also perform strongly on several culturally grounded Arabic benchmarks, including poetry, religion, cuisine, and dream interpretation, as well as in general tasks such as translation and summarization. We release the models in HuggingFace under a commercially permissive license. Jais 2 70B is also released as a chat app on the Web, iOS, and Android; it runs on Cerebras hardware, delivering up to 2,000 tokens per second, and enabling high-throughput Arabic-centric chat serving in our deployment setting. By uniting scale, linguistic diversity, cultural fidelity, openness, and speed, Jais 2 provides an open-weight foundation intended to support further research and development in Arabic-centric LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑