发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Borg框架,通过智能体驱动仿真器开发缩小LLM服务系统与仿真器的差距,其构建的扩展吞吐量误差更低,仿真速度显著优于现有仿真器。
AI 中文摘要
系统级仿真对于探索快速扩展的LLM服务系统设计空间至关重要,而实际部署这类系统成本高昂且往往不可行。然而,现代LLM服务的发展速度远超人工驱动仿真器开发的跟踪速度,从智能体工作流到拆分式服务的新兴工作负载与机制,已不再适配现有仿真器所假设的整体式仿真流程。因此,每一种新机制都需要进行侵入式重写,导致已部署服务系统与对其建模的仿真器之间的开发差距不断扩大。为缩小这一差距,本文提出Borg,一个实现智能体驱动仿真器开发的框架。Borg引入了可组合的仿真器基础设施,统一表达完整的服务工作流(包括协调工作流的控制决策),并将其在Borg仿真器中实现为统一动态图。合成智能体(Synthesizer agent)作为受控编码智能体,会在仿真器特定的约束和保真度验证下,将自然语言特征请求转换到该抽象层,从而演进一个共享仿真器,而非为每个特征构建新仿真器。在相同的编码智能体和受控条件下,基于Borg构建的扩展与基于vLLM的真实系统相比,平均吞吐量误差为2.51%,而基于现有仿真器构建的扩展误差为6.03%;在相同工作负载下,Borg的仿真速度分别比两个最先进的仿真器LLMServingSim2.0和Vidur快达284.96倍和23.19倍。
英文摘要
System-level simulation is an essential tool for exploring the rapidly expanding design space of LLM serving systems, where real deployments remain costly and often infeasible. However, modern LLM serving now evolves faster than human-driven simulator development can track, and emerging workloads and mechanisms, from agentic workflows to disaggregated serving, no longer fit the monolithic simulation pipeline that existing simulators assume. Each new mechanism therefore demands an invasive rewrite, leaving a widening development gap between deployed serving systems and the simulators that model them. To close this gap, we present Simthesizer, a framework that realizes agent-driven simulator development. Simthesizer introduces a composable simulator infrastructure that uniformly expresses the complete serving workflow, including the control decisions that coordinate it, and realizes it as a unified dynamic graph in Simthesizer simulator. Synthesizer agent, a harnessed coding agent, then lowers natural-language feature requests onto this abstraction under simulator-specific guardrails and fidelity validation, evolving one shared simulator instead of building a new one for every feature. Under the same coding agent and harnesses, extensions built on Simthesizer follow a vLLM-based real system with 2.51% average throughput error, versus 6.03% for extensions built on existing simulators. On identical workloads, Simthesizer also simulates up to 284.96x and 23.19x faster than two state-of-the-art simulators, LLMServingSim2.0 and Vidur, respectively.