发表机构
Dongguk University(东国大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对编译期静态LLM服务中工具等待与重新到达的成本问题,识别三种机制,提出模拟器选择配置,在并发8时最大批次调整可降成本8.25%,综合调整降9.72%。
AI 中文摘要
在智能体LLM服务中,会话调用外部工具,等待其返回,然后重新到达以继续推理。静态编译的NPU服务可以在编译时固定批次桶集合、最大批次大小以及KV缓存槽数量。我们将这种环境定义为编译期静态服务基板,并分析工具等待与重新到达在其中产生的执行期成本。在单个LLM实例上,我们运行遵循实测工具等待时间分布的合成工作负载,并在相同输入上,将基线配置(设置{1,2,4,8}、8和8)与改变其中部分设置的配置进行比较。由于不同配置下重新到达时间不同,我们构建了一个按时间顺序重放请求处理的模拟器,在2,077个配置中选择预测成本最低的候选配置,并在新输入上验证。我们识别出三种机制:离散批次对齐、KV缓存存续和预填充干扰。在并发度为6时,当批次桶集合中缺少某些桶,工具等待降低了填充率(从0.235降至0.120),但解码执行时间增加了1.51倍,因此填充率本身并不能指示成本。在并发度为8的新输入上,基线在24次重新到达中有9次重用KV,仅增大最大批次大小就削减了8.25%的执行成本,而所选配置(同时调整了批次桶集合)削减了9.72%。在18次重新到达中已有17次被重用的情况下,效果仅为0.59%。因此,编译期配置应通过诊断KV重用损失及其导致的执行变化来选择。
英文摘要
In agentic LLM services, a session calls an external tool, waits for it, and re-arrives to continue inference. Statically compiled NPU serving can fix the batch bucket set, the maximum batch size, and the number of KV cache slots at compile time. We define such an environment as a compile-time-static serving substrate and analyze the execution-time cost that tool waiting and re-arrival incur in it. On a single LLM instance, we run synthetic workloads following a measured tool waiting time distribution and compare, on the same inputs, a baseline configuration with settings {1, 2, 4, 8}, 8, and 8 against configurations that change some of them. Because re-arrival times differ across configurations, we build a simulator that replays request processing in time order, select the candidate with the lowest predicted cost among 2,077 configurations, and validate it on new inputs. We identify three mechanisms: discrete batch alignment, KV cache survival, and prefill interference. At a concurrency of 6, absent from the bucket set, tool waiting lowered the padding ratio (0.235 to 0.120) yet increased decode execution time 1.51-fold, so padding alone did not indicate cost. On new inputs at a concurrency of 8, where the baseline reused KV in 9 of 24 re-arrivals, enlarging the maximum batch size alone cut execution cost by 8.25%, and the selected configuration, which also adjusted the bucket set, by 9.72%. Where 17 of 18 re-arrivals were already reused, the effect was 0.59%. Compile-time configurations should thus be selected by diagnosing KV reuse loss and the resulting change in execution.
Comments24 pages, including 7 pages of supplementary material