arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

实际运行的是什么:Apple神经引擎上语言模型部署与解码速度的测量研究

How Weight Encoding Affects Language Model Placement and Performance on the Apple Neural Engine

Shahir M A

arXiv 2608.22110首次发表:更新:

AI 中文总结

本研究通过三项测量探究语言模型在Apple神经引擎的部署与解码速度,发现部署由计算表达而非内容决定,权重编码影响加速器使用,提出先选编码再分配参数的设计流程。

AI 中文摘要

本研究探究了哪些内容会将语言模型部署到Apple神经引擎(ANE)上运行,以及是什么因素决定了模型在ANE上的运行速度,通过三项测量工作给出答案:首先,在保持计算内容不变的前提下,对LLM原语按计算表达方式进行变化,构建了64种形态的矩阵,记录了各操作的设备支持情况;其次,训练了不同规模和精度的匹配模型,其中量化检查点的结构与对应的fp16模型完全一致,因此所有部署测量均针对真实训练后的产物;最后,在推理过程中读取ANE的内存控制器字节计数器,以此确定实际运行的内容而非编译器意图。每一项核心结论都至少由这三项测量路径中的两项提供支撑。研究发现,部署位置由计算的表达方式而非计算内容决定:融合RMSNorm完全可在ANE上运行,而其算术等效的分解形式仅能在CPU上运行;权重编码决定了加速器的使用情况:CoreML会将一个2585万参数、以卷积为主的fp16模型全部分配给CPU(计数器确认无字节通过引擎),而同一图的int8或2-bit版本则有约83%的驻留率,运行速度提升1.8至2.2倍,而一个规模更小的2229万参数的全注意力fp16模型的驻留率达98.9%。解码成本为每token的流送字节数,在fp16、int8和2-bit精度下均为标称编码宽度的约0.77倍。测量得到的最小且最快的模型是三值模型,在规模匹配的情况下,算子组合几乎不会改变这两个指标:所有驻留的2500万参数三值模型均在10.0至10.8 MB、0.62至0.64 ms/token的范围内。核心对比对象是半注意力2500万参数三值模型(10.5 MB,0.63 ms)和5000万参数三值模型(16.8 MB,0.86 ms),与本研究初始的以卷积为主的fp16设计相比,体积缩小9.8倍和6.1倍,速度提升3.0倍和2.2倍。基于这些测量结果,本研究提出了一个设计流程:先选择编码方式,再将字节预算分配给参数。

英文摘要

Weight compression can alter accelerator placement as well as memory traffic, complicating the interpretation of inference speedups. We investigate this interaction on the Apple Neural Engine through the public Core ML deployment path. Five independently trained language-model checkpoints span two architectures and dense fp16, int8, and ternary weights encoded with two-bit lookup tables. We combine compiler device plans, synchronized memory-controller measurements, and compute-unit exclusion controls for a fixed single-token forward workload. On an M1, the smaller fp16 export executes on the CPU despite permitting ANE execution, whereas its compressed counterparts exhibit ANE activity. The int8 export reduces warm forward latency by a factor of 1.9. The larger fp16 export also uses the ANE, indicating that dense encoding alone does not determine placement. Separate M3 energy measurements support the same direction of change. These results establish encoding-dependent placement in the measured deployment stack and show that compression comparisons require joint measurement of backend selection, latency, and traffic.

Comments10 pages, 2 figures, 7 tables. v2: corrected byte-traffic accounting, withdrew the 0.77 nominal-byte factor, and attributed normalization eligibility to dtype rather than formulation; title and abstract revised accordingly. See Section 8

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑