InstantInfer:使用通信有限自动机实现快速大语言模型冷启动
InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata
浏览论文内容
中文总结 AI 辅助
研究大语言模型推理服务冷启动低效问题,提出通信有限自动机抽象及编程框架,经证明其正确性后应用于vLLM相关环节形成InstantInfer,大幅加速冷启动且在多场景表现鲁棒。
中文摘要 AI 辅助
大语言模型推理服务中的冷启动严重影响用户体验,由于顺序初始化和复杂软件组件发出的大量细粒度I/O请求,冷启动效率仍然很低。尽管重构程序可以带来诸如并发执行和I/O合并等优势,但这种方法在处理大量异构组件时容易出错且存在正确性风险。我们提出通信有限自动机(CFA)抽象来系统地分析跨组件优化机会,并设计一个编程框架来实现基于CFA的组件程序重构。该框架保留了原始顺序程序结构,同时实现安全的并发组件执行。我们证明了程序重构的正确性。我们将CFA抽象和框架应用于vLLM中的进程树创建、张量加载和模型切换重构,形成了一个名为InstantInfer的新冷启动系统。大量实验表明,InstantInfer显著加速了大语言模型冷启动(加速高达7.2倍),并在不同的GPU、工作负载和规模上表现出鲁棒性。
英文摘要
Cold starts in large language model (LLM) inference services significantly affect user experience, yet they remain inefficient due to sequential initialization and a massive number of fine-grained I/O requests issued by complex software components. Although refactoring the program can yield advantages such as concurrent execution and I/O merging, this approach is error-prone and carries correctness risks when dealing with massive, heterogeneous components. We propose the Communicating Finite Automata (CFA) abstraction to systematically analyze cross-component optimization opportunities, and design a programming framework to enable CFA-based component program refactoring. This framework preserves the original sequential program structure while enabling safe concurrent component execution. We prove the correctness of the program refactoring. We apply the CFA abstraction and framework to refactor process tree creation, tensor loading, and model switching in vLLM, forming a new cold-start system named InstantInfer. Extensive experiments demonstrate that InstantInfer substantially accelerates LLM cold starts (achieving up to 7.2 times speedup) and exhibits robustness across diverse GPUs, workloads, and scales.