arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Vosti:指定、实现和验证确定性LLM推理

Vosti: Specifying, Implementing, and Verifying Deterministic LLM Inference

Jianxing Qin, Alexander Du, Danfeng Zhang, Matthew Lentz, Danyang Zhuo

arXiv 2609.38981首次发表:更新:

发表机构

Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Vosti是一个针对确定性LLM推理的推理引擎,通过形式化规范、内核选择独立性和KV缓存绑定,在多种执行变化下实现逐位相同输出,性能与vLLM相当。

AI 中文摘要

LLM推理系统可能改变批次组成、提示分块、预填充/解码执行以及KV缓存的重用、驱逐或重新计算。这些优化不应影响系统输出。生产系统,包括vLLM的批次不变模式和SGLang的确定性模式,旨在实现这一目标,但缺乏正式的系统级规范。我们形式化了确定性LLM推理:在固定的模型和部署配置下,具有相同提示和初始采样器状态的请求,在不同执行中,在对应的输出位置产生逐位相同的logits。我们的测试发现,这些生产模式在某些执行变化下会产生不同的logits。为解决这一局限,我们提出Vosti,一个针对此规范设计并验证的推理引擎。Vosti独立于运行时引擎状态选择内核,并将缓存的KV值绑定到其逻辑令牌前缀。其证明在引擎/GPU内核边界处分解:一个Verus归纳证明确立了调度和分页、前缀共享的KV缓存保持输出logits,而一个Triton分析器证明了所选内核输出在批次、查询长度和分页KV缓存布局中逐位相等。Vosti在所有测试的执行变化中产生逐位相同的logits,并在解码密集型工作负载上实现与vLLM的批次不变模式相当的性能,同时提供更强、正式验证的确定性保证。

英文摘要

LLM inference systems may vary batch composition, prompt chunking, prefill/decode execution, and KV-cache reuse, eviction, or recomputation. These optimizations should not affect system outputs. Production systems, including vLLM's batch-invariant mode and SGLang's deterministic mode, target this goal but lack a formal system-level specification. We formalize deterministic LLM inference: under a fixed model and deployment configuration, requests with the same prompt and initial sampler state produce bitwise-identical logits at corresponding output positions across executions. Our tests find that these production modes produce different logits under some execution variations. To address this limitation, we present Vosti, an inference engine designed and verified against this specification. Vosti chooses kernels independently of runtime engine state and ties cached KV values to their logical token prefixes. Its proof decomposes at the engine/GPU kernel boundary: a Verus inductive proof establishes that scheduling and the paged, prefix-sharing KV-cache preserve output logits, while a Triton analyzer proves bitwise-equal selected kernel outputs across batches, query lengths, and paged KV-cache layouts. Vosti produces bitwise-identical logits across every tested execution variation and achieves performance comparable to vLLM's batch-invariant mode on decode-heavy workloads, while providing a stronger, formally verified determinism guarantee.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑