面向可持续分布式LLM推理:面向能源、碳和缓存感知的llm-d控制平面的系统综合与研究议程
Toward Sustainable Distributed LLM Inference: A Systems Synthesis and Research Agenda for an Energy-, Carbon-, and Cache-Aware llm-d Control Plane
浏览论文内容
中文总结 AI 辅助
本文综合分布式LLM推理的可持续性研究,提出面向llm-d控制平面的能源、碳和缓存感知设计模式与评估框架,将可持续推理视为跨多维度控制问题。
中文摘要 AI 辅助
大语言模型(LLM)的可持续性日益成为一个服务系统问题,而不仅仅是训练问题。在生产环境中,能源和碳排放的影响不仅取决于模型大小:工作负载形态、批处理、键值(KV)缓存重用、预填充/解码放置、模型和加速器选择、电源状态、地理碳强度以及服务级别目标(SLO)都至关重要。最近的系统论文分别研究了这些因素中的许多方面。本文将这些结果联系起来,并提出一个实际的工程问题:当决策点是像llm-d这样的分布式推理控制平面时,这些结果意味着什么?这里的贡献是综合,而不是一组新的基准测试结果。报告的性能、能源、碳和成本改进仍然是所引用论文和系统的结果。我将文献归纳为反复出现的设计模式,并利用这些模式为llm-d勾勒出一个可持续推理控制平面(SICP)。所提出的控制平面将在路由和扩展时考虑延迟、能源、碳、缓存重用、服务成本和质量,同时将TTFT/TPOT SLO作为硬约束。我还概述了一个基于每焦耳和每克CO2e的SLO满足的goodput评估框架,以及一个可重复的实验计划。从文献联系中得到的主要观察是,可持续的LLM推理不太可能来自一个“绿色”模型或一个加速器;它更自然地被视为跨模型、阶段、缓存、硬件、副本、区域和时间的控制问题。
英文摘要
Large language model (LLM) sustainability is increasingly a serving-systems problem, not only a training problem. In production, energy and carbon impact depend on more than model size: workload shape, batching, key-value (KV) cache reuse, prefill/decode placement, model and accelerator choice, power state, geographic carbon intensity, and service-level objectives (SLOs) all matter. Recent systems papers study many of these factors separately. This paper connects those results and asks a practical engineering question: what do they imply when the decision point is a distributed inference control plane such as llm-d? The contribution here is synthesis, not a new set of benchmark results. Reported performance, energy, carbon, and cost improvements remain the results of the cited papers and systems. I group the literature into recurring design patterns and use those patterns to sketch a Sustainable Inference Control Plane (SICP) for llm-d. The proposed control plane would consider latency, energy, carbon, cache reuse, serving cost, and quality when routing and scaling, while keeping TTFT/TPOT SLOs as hard constraints. I also outline an evaluation framework based on SLO-satisfied goodput per joule and per gram CO2e, together with a reproducible experimental plan. The main observation from connecting the literature is that sustainable LLM inference is unlikely to come from one "green" model or one accelerator; it is more naturally treated as a control problem across model, phase, cache, hardware, replica, region, and time.