一次编写,随处运行:用于形状安全且框架无关的大语言模型架构的Axon领域特定语言
Write Once, Run Everywhere: The Axon DSL for Shape-Safe and Framework-Agnostic LLM Architectures
浏览论文内容
中文总结 AI 辅助
针对大语言模型架构的可移植性与效率瓶颈,提出强类型领域特定语言Axon,可编译为多框架实现,在467项基准测试中实现显著加速,突破部署锁定问题。
中文摘要 AI 辅助
开源语言模型的整个生态系统实际上依赖于单一平台。如果这个平台明天被迫关闭会怎样?实现和维护高效的模型定义,并在不同的训练和推理范式之间进行转换,是一项资源密集型任务,严重限制了模型的效率和可移植性,阻碍了模型的扩展和部署。在此,我们提出Axon,一种具有类Haskell语法的强类型领域特定语言,它为大语言模型架构实现了一次编写、随处运行的范式。通过基于语言规范而非特定框架的愿景开展协作,Axon促进了开放合作,使研究人员能够实现高度专业化的架构,而无需放弃优化基础设施或接受部署锁定。Axon允许生成简洁、可审计的规范,这些规范可自动编译为主流框架的独立实现:PyTorch、带Triton的PyTorch、JAX、MLX和vLLM。在针对参数规模从1.35亿到320亿的模型开展的467项推理基准测试实验中,与来自Transformers的参考实现相比,Axon在PyTorch上实现了7%的中位数加速,在带Triton的PyTorch上实现了12%的中位数加速,在JAX上实现了91%的中位数加速,在MLX上实现了107%的中位数加速;当作为带有PagedAttention和KV缓存的原生vLLM架构部署时,Axon模型实现了比Transformers实现高58%的中位数加速。
英文摘要
The entire ecosystem of open-source language models effectively relies on a single platform. What if this platform was forced to shut down tomorrow? Implementing and maintaining efficient model definitions and translating them between different training and inference regimes is a resource-heavy task that severely limits model efficiency and portability, hindering both scaling and deployment. Here, we present Axon, a strongly typed domain-specific language with Haskell-like syntax, that enables a write-once, run everywhere paradigm for LLM architectures. By basing collaboration on a language specification rather than a specific framework's vision, Axon fosters open cooperation and empowers researchers to implement highly specialized architectures without giving up optimization infrastructure or accepting deployment lock-in. Axon allows for concise, auditable specifications that can be automatically compiled to standalone implementations for leading frameworks: PyTorch, PyTorch with Triton, JAX, MLX and vLLM. In 467 inference benchmarking experiments on models ranging from 135M to 32B parameters, we demonstrate median speedups of 7% on PyTorch, 12% on PyTorch with Triton, 91% on JAX, and 107% on MLX, compared to the reference implementations from Transformers. When deployed as native vLLM architectures with PagedAttention and KV-cache, Axon models achieve a 58% median speedup over Transformers implementations.