AI 中文总结
本研究针对WebGPU用于LLM推理时的调度开销问题,提出顺序调度测量方法,发现批大小为1时调度开销是瓶颈,核心贡献为明确该优化方向并指出调度摊销的改进路径。
AI 中文摘要
大语言模型(LLM)被部署在从互联网浏览器到边缘设备的多种环境中,WebGPU是一种现代跨平台标准。基于浏览器的LLM推理引擎已大量涌现,但WebGPU的逐操作调度开销仍未得到充分表征。在本研究中,我们提出了一种顺序调度测量方法,结果表明,单纯的逐操作测量会因将调度与同步混淆而高估每次调度的成本。使用该方法,我们测量了每次调度的成本,发现其与所用数据类型无关。我们还表明,当批大小为1时,瓶颈在于调度开销而非内核质量,且调度计数是造成该问题的原因。因此,我们得出结论:当批大小为1时,在WebGPU中优化LLM推理的有效方法是减少调度计数。我们的发现指出,在推理引擎和WebGPU规范中进行调度摊销,是实现实用基于浏览器的推理的一条途径。
英文摘要
Large Language Models are deployed to multiple types of environments, from internet browsers to edge devices, and WebGPU serves as a modern cross-platform standard. The engines for browser-based LLM inference have proliferated, yet the overhead of WebGPU per-operation dispatch remains poorly characterized. In this work, we introduce a sequential-dispatch measurement method and show that naive single-operation measurements overestimate per-dispatch cost by conflating dispatch with synchronization. Using our method, we measure the per-dispatch cost and show that it is independent of data type used. We show that the dispatch overhead, not kernel quality, is the bottleneck at batch size 1, and isolate the dispatch count as the cause. Therefore, we conclude that at batch size 1, the effective approach to LLM inference optimization in WebGPU is reducing dispatch count. Our findings point to dispatch amortization, in the inference engines and in the WebGPU specification, as a path to practical browser-based inference.
CommentsAccepted to The Fourth UK AI Conference 2026 as a full paper, to be published in Proceedings of Machine Learning Research (PMLR)