arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35188cs.AIcs.PF

Token之下:GPU加速LLM推理中多Token预测的性能工程研究

Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference

Suwesh Prasad Sah

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在NVIDIA A10G GPU上对比了双Token多Token预测与自回归解码,发现MTP通过摊销效应将吞吐量提升1.91至2.19倍,并减少重复CUDA图执行,从而改善推理性能。

中文摘要 AI 辅助

自回归大语言模型推理会反复调用目标模型以每次生成一个Token,这使得生成过程对GPU内存移动和顺序执行非常敏感。本研究在受控的单请求部署环境中,于NVIDIA A10G GPU上评估了双Token多Token预测(MTP)与自回归解码的性能对比。一个包含360个请求的基准测试覆盖了纯文本、推理密集型及工具调用工作负载,同时使用运行时遥测、Nsight Systems、PyTorch Profiler以及选定的Nsight Compute测量来解释观察到的性能差异。MTP在所有提示词上将输出吞吐量提高了1.91倍至2.19倍,并将首个输出时间缩短了10.0%至14.2%。平均接受长度的中位数范围为每次验证迭代2.370至2.595个Token。性能分析显示,MTP引入了更长且更复杂的执行路径,包括提议、采样、注意力、聚合和归约操作。然而,每个生成Token所需的选定重复CUDA图执行次数减少了56.4%至78.1%。占主导地位的MTP GEMM内核并不比占主导地位的自回归GEMV内核更快,且两者的选定实例均接近A10G内存带宽极限。这些结果表明,MTP通过摊销效应改善了推理性能:更大的Token进展减少了重复GPU执行的次数,足以抵消额外的推测执行成本。

英文摘要

Autoregressive large language model inference repeatedly invokes the target model to generate one token at a time, making generation sensitive to GPU memory movement and sequential execution. This study evaluates two-token multi-token prediction (MTP) against autoregressive decoding in a controlled single-request deployment on an NVIDIA A10G GPU. A 360-request benchmark covered plain-text, reasoning-intensive, and tool-calling workloads, while runtime telemetry, Nsight Systems, PyTorch Profiler, and selected Nsight Compute measurements were used to explain the observed performance. MTP increased output throughput by \(1.91\times\) to \(2.19\times\) across all prompts and reduced time to first output by 10.0--14.2\%. Median mean acceptance length ranged from 2.370 to 2.595 tokens per verification iteration. Profiling showed that MTP introduced a longer and more complex execution path, including proposal, sampling, attention, gathering, and reduction operations. However, it required 56.4--78.1\% fewer executions of the selected repeating CUDA Graph per generated token. The dominant MTP GEMM kernel was not faster than the dominant autoregressive GEMV kernel, and selected instances of both approached the A10G memory-bandwidth limit. These results show that MTP improved inference through amortization: greater token progress reduced repeated GPU execution sufficiently to outweigh the additional speculative-execution cost.

↑