arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoRun:填充对于确定性大语言模型(LLM)推理而言简单且高效

CoRun: Padding is Simple and Efficient for Deterministic LLM Inference

Shiju Zhao, Jiacheng Yang, Qihang Chen, Junhao Hu, Jiaqi Zheng, Guihai Chen, Xusheng Chen

arXiv 2608.14376首次发表:更新:

AI 中文总结

CoRun利用LLM内核的位置不变性,通过独立预填充和固定形状批量解码的调度系统,在保证确定性推理的同时提升了LLM的吞吐量并降低了延迟。

AI 中文摘要

尽管采样参数和随机种子固定,但大语言模型(LLM)推理仍存在输出不一致问题,这会损害模型评估、强化学习等下游任务。这种非确定性的一个主要来源是依赖批次的GPU执行:动态输入形状会改变内核分块和浮点归约顺序。现有系统通过批次不变内核解决该问题,但这些内核限制了优化分块和拆分归约,导致延迟增加2倍以上,服务吞吐量最多降低74%。本文发现,虽然大多数内核并非批次不变,但它们具有位置不变性。利用这一特性,我们提出CoRun,一种基于调度的系统,无需批次不变性即可实现确定性推理。CoRun采用独立的预填充和固定形状的批量解码分别处理LLM推理的两个阶段,利用CUDA图实现高效执行并简化实现。在Qwen、DeepSeek等不同架构的LLM上进行的实验表明,CoRun在保证确定性的同时,相比批次不变方法将吞吐量提升了15%-324%,平均首token时间降低51.8%,每输出token时间降低48.6%。

英文摘要

Despite fixed sampling parameters and random seeds, Large Language Model (LLM) inference exhibits output inconsistency, which undermines downstream tasks such as model evaluation and reinforcement learning. A major source of this nondeterminism is batch-dependent GPU execution: dynamic input shapes change kernel tiling and floating-point reduction orders. Existing systems address this problem with batch-invariant kernels, but these kernels restrict optimized tiling and split reductions, increasing more than 2$\times$ latency and reducing serving throughput by up to 74 %. This paper observes that although most kernels are not batch-invariant, they are position-invariant. Leveraging this property, we present CoRun, a scheduling-based system that achieves deterministic inference without requiring batch invariance. CoRun employs isolated prefill and fixed-shape batched decode to handle the two stages of LLM inference, respectively, leveraging CUDA graphs for efficient execution and simplified implementation. Experiments on LLMs with diverse architectures, including Qwen and DeepSeek, show that CoRun ensures determinism while improving throughput by 15-324 % over batch-invariant approaches, reducing time-to-first-token by 51.8 % and time-per-output-token by 48.6 % on average.

Comments13 pages, 14 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑