GGUF元数据对单序列llama.cpp吞吐量在三系统上的预测
GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems
浏览论文内容
中文总结 AI 辅助
本研究利用GGUF元数据和屋顶线模型预测llama.cpp单序列吞吐量,在多种硬件上实现较低误差,并发现效率系数不具普适性。
中文摘要 AI 辅助
我们利用基于屋顶线形状的预测器,结合在参考模型上拟合的量化特定比例因子,从GGUF元数据预测单序列模型吞吐量。评分队列包含来自两个Apple M4 Max系统和一个NVIDIA RTX 5080上的53个主机-文件配置的318个阶段-深度测量值。在主机特定的保留集(分别包含四个、五个和两个配置)上,使用活跃参数的解码模型获得了13.1%、14.4%和36.1%的平均绝对百分比误差(MAPE),而按总参数计费时,误差分别为49.4%、55.3%和51.9%。在其他两个系统上拟合的留一主机系数在测试集上产生了11.6%、16.8%和36.0%的MAPE。低比特模型阶梯改变了运行时堆栈之间的排序。P2预填充基线在测试集上的MAPE为18.7%、22.2%和108.2%。GGUF结构在所有三个系统上都有帮助,但拟合的效率并非普遍适用。
英文摘要
We predict single-sequence model throughput from GGUF metadata using roofline-shaped predictors with quantization-specific scale factors fitted on reference models. The scored cohort comprises 318 phase-depth measurements from 53 host-file configurations on two Apple M4 Max systems and an NVIDIA RTX 5080. On host-specific held-out sets of four, five, and two configurations, an active-parameter decode model obtains 13.1%, 14.4%, and 36.1% mean absolute percentage error (MAPE), versus 49.4%, 55.3%, and 51.9% when charging total parameters. Leave-one-host-out coefficients fitted on the other two systems yield 11.6%, 16.8%, and 36.0% test MAPE. A low-bit model ladder changes ordering across runtime stacks. The P2 prefill baseline gives 18.7%, 22.2%, and 108.2% test MAPE. GGUF structure helps on all three systems, but fitted efficiencies are not universal.
发表机构
- Northeastern University(东北大学)
- Sofia University(索菲亚大学)
- Indiana University(印第安纳大学)
- University of California, San Diego(加利福尼亚大学圣迭戈分校)
- Michigan State University(密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。