发表机构
Samsung Advanced Institute of Technology(三星先进技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ServeTwin是一个闭环模拟器,通过分析时间预测与有状态服务循环结合,无需硬件剖析即可高精度模拟分布式LLM服务,支持架构探索。
AI 中文摘要
在物理集群上评估分布式LLM服务设计成本高昂。然而,现有模拟器仅提供现实设计探索所需能力的子集:有状态闭环执行、无需对目标硬件进行性能剖析的时间预测,以及直接执行未经修改的服务基准测试。我们提出ServeTwin,一种闭环模拟器,将规格驱动的分析时间预测与有状态服务循环相结合。这种结合捕获了请求完成、调度、队列状态和KV缓存演化之间的反馈,包括预填充-解码分离和多轮智能体工作负载。ServeTwin通过iSTAGE避免目标硬件算子性能剖析,iSTAGE是一种分析轨迹生成器,从模型和引擎规格推导每次迭代的轨迹,同时表示非均匀批次和离散引擎机制,如CUDA图批次填充。它进一步将执行时间分解为组件拥有的吞吐量、调度器和运行时成本。这种归属允许不受影响的参数跨平台转移,并将重新校准限制在更改的硬件或软件组件上。ServeTwin实现了与vLLM兼容的接口,并运行未经修改的服务基准测试。针对真实部署,它以3.6%的平均误差重现了InferenceX的稳态吞吐量-交互性前沿,并以9.9%的误差预测了LMBenchmark的多轮性能,同时跟踪KV缓存演化。硅前扫描揭示,首选的HBM带宽-容量权衡在工作负载状态间反转,且增加并发将瓶颈从内存带宽转移到调度器和运行时开销。总之,这些能力使得在目标硬件可用之前对分布式LLM服务系统进行实际探索成为可能。
英文摘要
Evaluating distributed LLM serving designs on physical clusters is costly. Yet existing simulators provide only subsets of the capabilities needed for realistic design exploration: stateful closed-loop execution, timing prediction without profiling target hardware, and direct execution of unmodified serving benchmarks. We present ServeTwin, a closed-loop simulator that couples specification-driven analytical timing with a stateful serving loop. This coupling captures feedback among request completion, scheduling, queue state, and KV-cache evolution, including prefill-decode disaggregation and multi-turn agent workloads. ServeTwin avoids target-hardware operator profiling through iSTAGE, an analytical trace generator that derives per-iteration traces from model and engine specifications while representing ragged batches and discrete engine mechanisms such as CUDA-graph batch padding. It further decomposes execution time into component-owned throughput, scheduler, and runtime costs. This ownership lets unaffected parameters transfer across platforms and confines recalibration to changed hardware or software components. ServeTwin implements a vLLM-compatible interface and runs unmodified serving benchmarks. Against real deployments, it reproduces InferenceX's steady-state throughput-interactivity frontier with a 3.6% mean error and predicts LMBenchmark's multi-turn performance with a 9.9% error while tracking KV-cache evolution. Pre-silicon sweeps reveal that the preferred HBM bandwidth-capacity tradeoff reverses across workload states and that increasing concurrency shifts the bottleneck from memory bandwidth to scheduler and runtime overheads. Together, these capabilities enable practical exploration of distributed LLM serving systems before target hardware is available.