arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁为KV缓存买单?在Kubernetes和LLM提供商账单中归因共享AI推理支出

Who Pays for the KV Cache? Attributing Shared AI Inference Spend Across Kubernetes and LLM Provider Bills

Timothy Urista

arXiv 2609.24991首次发表:更新:

AI 中文总结

本文提出unalloc工具,统一Kubernetes与LLM提供商账单,揭示共享推理支出归因失效问题,并实验证明计量规则显著影响KV缓存成本分配。

AI 中文摘要

组织通过不连贯的账本为AI付费:用于自托管推理的Kubernetes分配、网关日志以及API提供商按token计费的账单。我们提出了unalloc,一个开源工具,它将OpenCost、LiteLLM、OpenAI和Anthropic的成本数据合并为一个精确的账本,并报告无所有者的支出份额,并使用它来研究这些系统之间接缝处归因失效的地方。五个案例研究运行真实推理或模拟推理:一个具有分页KV内存和前缀缓存的vLLM风格服务模拟器;一个PyTorch transformer,使用真实KV缓存服务多租户轨迹;在此HTTP URL上的张量和流水线并行推理;针对模拟提供商API的未修改CLI;以及四个下游用例。在接缝处,在一个构造的多Pod部署场景中——一个月的合成OpenCost分配,而非观察到的计费数据——仅在LeaderWorkerSet领导者Pod上设置所有者标签,导致该部署的GPU账单中有66%无人认领,而自然回退键将其中61%分配给Helm图表名称,而标题性的未分配份额降至4%;启用每个来源会使所有网关支出重复计算;读取计费API的一页报告了四分之一的支出。在共享推理服务器内部,计量规则决定了谁付费:在运行vLLM的NVIDIA H100上,token计量器在每次测试负载下将检索密集型租户的账单份额比等时间份额计量器高出12-14个百分点,而GPU利用率在配置负载为每秒2到16个请求(每秒3.7到26.9个已完成请求;配置速率仅计算会话初始到达)时读取为97-99%,功耗随负载变化。两种计量器都不是地面真相;我们将这些结果与最近基于Shapley的能量归因进行对比。代码、原始数据、捕获的证据、图表和论文均可从存储库重新生成。

英文摘要

Organizations pay for AI through disconnected ledgers: Kubernetes allocations for self-hosted inference, gateway logs, and per-token bills from API providers. We present unalloc, an open-source tool that joins OpenCost, LiteLLM, OpenAI and Anthropic cost data into one exact ledger and reports the share of spend with no owner, and use it to study where attribution breaks at the seams between these systems. Five case studies run inference for real or simulate it: a vLLM-style serving simulator with paged KV memory and prefix caching; a PyTorch transformer serving a multi-tenant trace with a real KV cache; tensor- and pipeline-parallel inference on torch.distributed; the unmodified CLI against mock provider APIs; and four downstream use cases. At the seams, in a constructed multi-pod deployment scenario -- one month of synthetic OpenCost allocations, not observed billing data -- owner labels set only on LeaderWorkerSet leader pods leave 66% of that deployment's GPU bill unowned, and the natural fallback key assigns 61% of it to a Helm chart name while the headline unallocated share falls to 4%; enabling every source double counts all gateway spend; and reading one page of a billing API reports a quarter of spend. Inside a shared inference server the metering rule decides who pays: on an NVIDIA H100 running vLLM, a token meter assigns a retrieval-heavy tenant 12-14 percentage points more of the bill than an equal time-share meter at every load tested, while GPU utilization reads 97-99% across configured loads of 2 to 16 requests per second (3.7 to 26.9 completed requests per second; the configured rate counts session-initial arrivals only) and power draw tracks load. Neither meter is a ground truth; we position these results against recent Shapley-based energy attribution. Code, raw data, captured evidence, figures and the paper regenerate from the repository.

Comments14 pages, 9 figures, 4 tables. Includes a validation run with vLLM on an NVIDIA H100. Code, data and reproduction scripts: https://github.com/timurista/unalloc. Software: doi:10.5281/zenodo.22761012. Use of generative AI is disclosed in the paper ("Use of AI tools")

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑