发表机构
Illinois Institute of Technology; University of Wisconsin–Milwaukee(伊利诺伊理工学院; 威斯康星大学密尔沃基分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PrefixBench-H100提出一个可复现基准,在H100上系统表征前缀复用对LLM服务性能的影响,识别其收益与受限条件,并揭示跨运行时差异源于调度层。
AI 中文摘要
重复的提示前缀在LLM服务工作负载中日益普遍,出现在系统提示、模板化的检索增强生成流水线、智能体框架和多轮对话中。诸如vLLM和TensorRT-LLM等现代推理运行时提供了跨请求复用先前计算的KV缓存状态的机制,但尚不清楚在当代加速器上前缀复用何时能实质性提升服务性能,以及其收益何时受到调度、缓存粒度、并发性或内存压力的限制。本文提出了PrefixBench-H100,一个可复现的基准测试和测量框架,用于表征在单个NVIDIA H100上的前缀复用。PrefixBench-H100将受控的合成轨迹与聊天风格和检索风格的工作负载相结合,并在匹配的工作负载条件下评估两个广泛使用的LLM服务运行时。该基准测试变化共享前缀长度、后缀多样性、请求到达模式、并发性、输出长度和缓存配置,同时收集首令牌时间、令牌间延迟、端到端延迟、吞吐量、缓存命中统计、GPU内存使用情况以及选定的性能剖析轨迹。PrefixBench-H100的目标不是引入新的缓存算法,而是揭示H100级LLM服务中前缀复用的实际运行范围。该研究识别了前缀复用带来显著首令牌延迟降低的机制,以及缓存压力侵蚀这些收益的机制,同时表明缓存有效性本身在很大程度上对并发性和输出长度不敏感;剩余的跨运行时差异出现在缓存之上,即调度层。
英文摘要
Repeated prompt prefixes are increasingly common in LLM serving workloads, appearing in system prompts, templated retrieval-augmented generation pipelines, agent frameworks, and multi-turn conversations. Modern inference runtimes such as vLLM and TensorRT-LLM provide mechanisms for reusing previously computed KV-cache state across requests, yet it remains unclear when prefix reuse materially improves serving performance on contemporary accelerators and when its benefits are limited by scheduling, cache granularity, concurrency, or memory pressure. This paper presents PrefixBench-H100, a reproducible benchmark and measurement framework for characterizing prefix reuse on a single NVIDIA H100. PrefixBench-H100 combines controlled synthetic traces with chat-style and retrieval-style workloads, and evaluates two widely used LLM serving runtimes under matched workload conditions. The benchmark varies shared-prefix length, suffix diversity, request arrival pattern, concurrency, output length, and cache configuration, while collecting time-to-first-token, inter-token latency, end-to-end latency, throughput, cache-hit statistics, GPU memory usage, and selected profiling traces. The goal of PrefixBench-H100 is not to introduce a new caching algorithm, but to expose the practical operating envelope of prefix reuse for H100-class LLM serving. The study identifies the regime where prefix reuse provides substantial first-token latency reductions and the regime where cache pressure erodes them, while showing that cache effectiveness itself is largely insensitive to concurrency and output length; the cross-runtime differences that remain arise above the cache, in the scheduling layer.