arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过匹配的 FP16 中间层分解运行时、内核和量化加速:基于四块 NVIDIA RTX A5000 GPU 的硬件条件案例研究

Decomposing Runtime, Kernel, and Quantization Speedups via a Matched FP16 Intermediate: A Hardware-Conditioned Case Study on Four NVIDIA RTX A5000 GPUs

Weijia Han, Lisha Qu

arXiv 2607.11368首次发表:更新:

发表机构

University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究通过匹配的 FP16 中间层分解运行时、内核和量化加速,以四块 NVIDIA RTX A5000 GPU 为案例,分析加速各部分占比、分片影响因素等,得出运行时占主要加速比例,量化可扩展并发用户,还探讨了实例选择及有效性威胁。

AI 中文摘要

报告的量化内核服务加速通常将权重格式、内核和推理运行时捆绑为一个数字。我们对单个主机上通过 NVLink 桥接的四台 24 GiB 的 NVIDIA RTX A5000 GPU 进行了归因研究。一个匹配的中间堆栈在没有量化内核的情况下保持更快的运行时,将全加速分解为运行时部分和内核与量化部分。在匹配的贪婪解码下,完整堆栈端到端达到 2.58 倍,运行时变化在对数尺度上占该增益的约三分之二;在三个相似模型系列中,内核和量化部分最多移动 1.5%。在所有四张卡上分片一个实例远低于翻倍:分析器跟踪将每个令牌短缺的约 80%归因于协调,并且同一硬件上的 NVLink 与 PCIe 控制在两条链路上显示出相似的实际带宽,表明链路带宽不是原因。运行一个分片实例还是几个独立实例取决于工作负载和模型,在较大模型上排名相反:较小模型根据工作负载在分片和多个实例之间划分,而较大模型在每个工作负载上更喜欢两个配对实例。量化将可持续并发用户扩展到可重现的半精度内存悬崖之后大约四倍。记录了两个堆栈之间采样模式和提示池的差异作为有效性威胁。

英文摘要

Reported serving speedups from quantized kernels typically bundle the weight format, the kernel, and the inference runtime into one number. We present an attribution study on four NVIDIA RTX A5000 GPUs, 24 GiB each, on a single host with NVLink-bridged pairs. A matched intermediate stack that keeps the faster runtime without the quantized kernel splits the full speedup into a runtime part and a kernel and quantization part. Under matched greedy decoding the full stack reaches $2.58\times$ end to end, with the runtime change accounting for about two thirds of that gain on a logarithmic scale; across three similar model families the kernel and quantization part moves by at most 1.5%. Sharding one instance across all four cards falls well below doubling: a profiler trace attributes about 80% of the per token shortfall to coordination, and an NVLink versus PCIe control on the same hardware shows similar realized bandwidth on both links, pointing away from link bandwidth as the cause. Whether to run one sharded instance or several independent ones depends on the workload and the model, with the ranking reversing on the larger model: the smaller model splits between sharding and multiple instances by workload, while the larger model favors two paired instances on every workload. Quantization extends sustainable concurrent users roughly four times past a reproducible half precision memory cliff. Differences in sampling mode and prompt pool between the two stacks are documented as threats to validity.

Comments36 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑