发表机构
TII; LightOn(技术创新研究院; LightOn公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文对超大型阿拉伯语语言模型Noor的碳足迹进行整体评估,涵盖数据、研发、预训练、推理及外生成本,发现推理与外生因素影响显著,并探讨减排路径。
AI 中文摘要
随着越来越大的语言模型变得日益普及,考虑其环境影响至关重要。以极端规模和资源消耗为特征,最近几代模型因其对计算资源的贪婪需求以及由此产生的显著碳足迹而受到批评。尽管机器学习论文中碳影响的报告已变得更加普遍,但这种报告通常仅限于严格用于训练的计算资源。在这项工作中,我们提出了对一个极端规模语言模型Noor足迹的整体评估。Noor是一个正在进行的项目,旨在开发最大的多任务阿拉伯语语言模型——拥有高达130亿参数——利用零样本泛化能力,通过自然语言指令实现广泛的 downstream 任务。我们评估了整个项目的总碳账单:从数据收集和存储成本开始,包括研发预算、预训练成本、未来服务估算,以及这一国际合作所必需的其他外生成本。值得注意的是,我们发现推理成本和外生因素可能对总预算产生显著影响。最后,我们讨论了减少极端规模模型碳足迹的途径。
英文摘要
As ever larger language models grow more ubiquitous, it is crucial to consider their environmental impact. Characterised by extreme size and resource use, recent generations of models have been criticised for their voracious appetite for compute, and thus significant carbon footprint. Although reporting of carbon impact has grown more common in machine learning papers, this reporting is usually limited to compute resources used strictly for training. In this work, we propose a holistic assessment of the footprint of an extreme-scale language model, Noor. Noor is an ongoing project aiming to develop the largest multi-task Arabic language models -- with up to 13B parameters -- leveraging zero-shot generalisation to enable a wide range of downstream tasks via natural language instructions. We assess the total carbon bill of the entire project: starting with data collection and storage costs, including research and development budgets, pretraining costs, future serving estimates, and other exogenous costs necessary for this international cooperation. Notably, we find that inference costs and exogenous factors can have a significant impact on total budget. Finally, we discuss pathways to reduce the carbon footprint of extreme-scale models.
Comments11 pages, 3 figures, 2 tables. Published in Proceedings of BigScience Episode #5 -- Workshop on Challenges & Perspectives in Creating Large Language Models (ACL 2022)
Journal refProceedings of BigScience Episode #5 -- Workshop on Challenges & Perspectives in Creating Large Language Models, pages 84-94, Association for Computational Linguistics, 2022
DOI:10.18653/v1/2022.bigscience-1.8