AI 中文总结
本综述梳理了各类面向大语言模型的AI硬件加速器架构,基于Transformer与Roofline分析指出内存是加速核心约束,对比各平台特性并展望了未来以内存为中心的异构系统发展方向。
AI 中文摘要
大语言模型(LLMs)对用于其训练和推理的硬件提出了前所未有的且仍在不断增长的需求。本综述全面梳理了面向LLMs的AI硬件加速器的全貌,涵盖云与边缘部署场景下的通用GPU、TPU、Trainium、Groq、Cerebras等定制ASIC、可重构FPGA、存内计算与近存计算架构,以及新兴的神经形态与光子计算方案。以Transformer的计算结构和Roofline分析为通用框架,研究表明,LLM加速的决定性约束并非算术性能,而是内存:自回归解码阶段受带宽限制,键值缓存的规模可与模型权重相当,数据移动占主导能耗。通过从计算、内存、能耗、可编程性和可扩展性维度对比各平台,发现不存在适用于所有工作负载的最优架构:GPU仍是灵活的默认选择和训练主力;面向特定领域的ASIC在稳定、高吞吐量工作负载的规模化部署中占据优势;存内计算是应对内存墙的近期最具前景的方案,正作为异构补充进入系统;而神经形态与光子计算虽具潜力,但尚未达到前沿规模的量产就绪状态。未来进展依赖硬件-算法协同设计以及以内存为中心的异构系统:对于大语言模型而言,内存系统已成为计算机的核心。
英文摘要
Large language models (LLMs) place unprecedented and still-growing demands on the hardware that trains and serves them. This review surveys the full landscape of AI hardware accelerators for LLMs, including general-purpose GPUs, custom ASICs such as TPUs, Trainium, Groq, and Cerebras, reconfigurable FPGAs, processing-in-memory and near-memory architectures, and emerging neuromorphic and photonic approaches across cloud and edge deployment. Using the transformer's computational structure and roofline analysis as a common framework, we show that the decisive constraint on LLM acceleration is not arithmetic but memory: the autoregressive decode phase is bandwidth-bound, the key-value cache can rival the model weights in size, and data movement dominates energy. Comparing platforms on compute, memory, energy, programmability, and scalability, we find that no single architecture is optimal across workloads: GPUs remain the flexible default and the workhorse of training; domain-specific ASICs win at scale for stable, high-volume workloads; processing-in-memory is the most promising near-term response to the memory wall, entering systems as a heterogeneous complement; and neuromorphic and photonic computing, while promising, are not yet production-ready at frontier scale. Future progress depends on hardware-algorithm co-design and heterogeneous, memory-centric systems: for large language models, the memory system has become the computer.
CommentsReview/survey article on AI hardware accelerators for large language models; compares GPUs, ASICs, FPGAs, processing-in-memory/near-memory, neuromorphic, and photonic architectures