arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种基于MLIR的大语言模型编译方法

An MLIR-Based Compilation Method for Large Language Models

Pengchao Hu, Zhibin Xin, Yifan Chen, Yangyang Zhou, Liang Wang, Xin Zhang

arXiv 2607.15865首次发表:更新:

发表机构

Sophgo Inc(算能科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型在专用硬件部署的挑战,提出基于MLIR的编译方法,用TopOp和TpuOp方言表示模型,分阶段编译,在TPU-MLIR编译器等项目实现,支持多种模型及量化部署形式。

AI 中文摘要

大语言模型已成为现代人工智能加速器的主要工作负载,但在专用硬件上部署仍面临两个核心挑战:如何将训练好的模型导入编译器友好的中间表示,以及如何在有限的片上内存下有效调度自回归推理循环。本文提出一种基于MLIR的大语言模型编译方法,通过TopOp和TpuOp两种算子方言进行说明。TopOp作为独立于源框架和目标芯片的高级图方言,负责表达模型语义;TpuOp作为目标硬件方言,承载与芯片相关的决策。模型先表示为TopOp,再逐层降低到TpuOp,最后生成可部署二进制文件。此外,每个Transformer层分为三个阶段进行静态编译。该方法已在TPU-MLIR编译器和LLM-TPU部署项目中实现,支持多种生成模型及量化和部署形式。

英文摘要

Large Language Models (LLMs) have become the dominant workload on modern AI accelerators, yet deploying them on specialized hardware still faces two core challenges: how to import a trained model into a compiler-friendly intermediate representation, and how to efficiently schedule the autoregressive inference loop under limited on-chip memory. This paper presents an MLIR (Multi-Level Intermediate Representation) based compilation method for large language models, illustrated using two dialects of operators, TopOp and TpuOp. TopOp serves as a high-level graph dialect that is independent of both the source framework and the target chip, and is responsible for expressing model semantics; TpuOp serves as the target hardware dialect, carrying chip-related decisions such as quantization, layer groups, and memory layout. A model is first represented as TopOp, then lowered layer by layer to TpuOp, and finally a deployable binary is generated. In addition, each Transformer layer is split into three stages for static compilation: prefill, prefill_kv (prefill with historical key-value cache), and decode, so as to accommodate the different computational characteristics of prompt-parallel processing and per-token generation. The method has been implemented in the TPU-MLIR compiler {https://github.com/sophgo/tpu-mlir} and the LLM-TPU deployment project {https://github.com/sophgo/LLM-TPU}, supporting a variety of generative models including the Qwen, Llama, InternVL, and MiniCPM-V series, as well as multiple quantization and deployment forms such as GPTQ, AWQ, and AutoRound.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑