AI 中文总结
本教程通过但丁《神曲》实例,结合架构分析与流量模型,揭示LLM训练中词语到网络比特流的转化过程及通信需求。
AI 中文摘要
大型语言模型(LLMs)将大量非结构化文本转化为用于语言生成和推理任务的语义模式。在其易用性背后隐藏着一个复杂的过程:词语变成token,token变成向量,向量最终产生比特流,流经高性能计算(HPC)系统。随着现代LLM增长到数十亿或数万亿参数,这一路径越来越多地跨越数千个互连的加速器,使得底层通信结构成为模型训练中一个关键且往往不透明的组成部分。本教程旨在引导读者了解从词语到网络流量的旅程,阐明语言如何在HPC训练系统中被转化为通信流。我们以但丁的《神曲》中的具体例子,说明模型架构、tokenization、嵌入和并行化策略如何塑造网络中交换数据的数量、结构和时序。我们将架构分析与解析流量模型和数值示例相结合,以表征LLM训练的通信需求。我们试图揭开词语如何穿越网络的神秘面纱,并为支持从文本到训练模型的旅程所需的网络需求提供实用见解。
英文摘要
Large Language Models (LLMs) transform vast collections of unstructured text into semantic patterns used for language generation and reasoning tasks. Behind their ease of use lies a complex process: words become tokens, tokens become vectors, and vectors ultimately give rise to streams of bits that flow through High-Performance Computing (HPC) systems. As modern LLMs grow to billions or trillions of parameters, this path increasingly unfolds across thousands of interconnected accelerators, making the underlying communication fabric a critical and often opaque component of model training. This tutorial aims to walk the reader through the journey from words to network traffic, shedding light on how language is translated into communication flows within HPC training systems. Using concrete examples from Dante's Divine Comedy, we illustrate how model architecture, tokenization, embeddings, and parallelization strategies shape the volume, structure, and timing of data exchanged across the network. We combine architectural analysis with analytical traffic models and numerical examples to characterize the communication requirements of LLM training. We try to demystify how words travel across the network and provide practical insights into the network requirements needed to support the journey from text to trained model.