arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19438cs.ARcs.AIcs.CLcs.DCcs.LGcs.PF

BaseRT:利用苹果M5神经加速器提升一流大语言模型推理性能

BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators

Fabian Waschkowski, Prabod Rathnayaka, Lukas Wesemann

首次发表
浏览论文内容

中文总结 AI 辅助

研究利用苹果M5神经加速器提升大语言模型推理性能,通过BaseRT无框架设计及添加特定内核,将计算密集型矩阵乘法经M5处理,大幅提升推理吞吐量及解码优势,确立了设备上大语言模型推理新性能上限。

中文摘要 AI 辅助

苹果M5一代引入了重新设计的GPU架构,每个核心都有专用神经加速器,即通过Metal~4张量API暴露的片上矩阵单元。我们展示了BaseRT,这是我们用于苹果硅上大语言模型的原生Metal推理运行时,利用这些单元将苹果硬件上的推理吞吐量大幅提升至超越[this http URL]和MLX。基于BaseRT无框架设计,我们添加了一系列手写的Metal~4张量核心内核(包括密集和专家混合GEMM以及闪存注意力预填充内核),将推理中计算密集型矩阵乘法通过M5神经加速器处理,而将内存密集型解码路径留在现有专用内核上。在苹果M5 Pro上,针对15种模型配置,BaseRT在提示处理吞吐量方面比[this http URL]高出6.4倍,比MLX高出3.9倍,在矩阵乘法占主导的专家混合模型上优势最大,同时在解码方面比[this http URL]领先1.75倍,比MLX领先1.33倍。这些结果为设备上的大语言模型推理确立了新的性能上限,并表明M5的张量核心是苹果硅上提示处理的决定性因素。BaseRT可在[this https URL]公开获取。

英文摘要

Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed through the Metal~4 tensor API. We show that BaseRT, our native Metal inference runtime for large language models on Apple Silicon, exploits these units to push inference throughput on Apple hardware substantially beyond both llama.cpp and MLX. Building on BaseRT's framework-free design, we add a family of hand-written Metal~4 tensor-core kernels (including dense and mixture-of-experts GEMM and flash-attention prefill kernels) that route the compute-bound matrix multiplications of inference through the M5 Neural Accelerators while leaving the memory-bound decode path on our existing specialised kernels. On an Apple M5 Pro, across fifteen model configurations spanning the Qwen3, Qwen3.5/3.6, Llama~3.2, and Gemma~4 families from sub-1B to 35B parameters, BaseRT delivers up to $6.4\times$ higher prompt-processing throughput than llama.cpp and $3.9\times$ higher than MLX, with the largest margins on the mixture-of-experts models where matrix multiplication dominates, while maintaining its lead on decode of up to $1.75\times$ over llama.cpp and $1.33\times$ over MLX. These results establish a new performance ceiling for on-device LLM inference and show that the M5's tensor cores are the decisive lever for prompt processing on Apple Silicon. BaseRT is publicly available at https://github.com/basecompute/baseRT.

发表机构

  • Base Compute

机构由 AI 辅助整理,请以论文原文为准。

↑