统一AI网关:联合模型路由与KV缓存管理的框架
Unified AI Gateway: A Framework for Joint Model Routing and KV Cache Management
- Huawei(华为)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多模型推理中的集成与KV缓存重建成本,提出统一AI网关框架,联合优化模型路由、缓存管理与计算放置,模拟显示TTFT加速最高13.28倍。
AI中文摘要:
大语言模型(LLM)推理日益跨越在规模、能力、价格和提供商方面各不相同的模型。这一转变给开发者带来了两类成本。其一是集成成本,即在众多模型中进行选择并在它们之间切换的成本。其二是推理成本,即当KV缓存不可用或与所选模型不兼容时,需要重建KV缓存所产生的成本。我们将统一AI网关定义并分析为一种边缘部署的AI流量枢纽的系统设置。它协调终端设备、边缘资源和云模型服务之间的模型路由、KV缓存管理和计算放置。在请求时,网关在任务质量、延迟、成本和资源约束下联合选择目标模型、执行站点和KV缓存操作。同时,后台缓存管理操作优化KV缓存的放置、复制、检索和生命周期决策,以服务于后续请求。我们综合了关于KV缓存复用、压缩、跨模型映射、分布式存储和传输的现有证据,并讨论了将这些能力集成到一个系统中仍面临的挑战。在八个典型工作负载配置下,我们的工作负载级分析模拟报告了TTFT加速1.25倍至13.28倍,输入成本效益1.20倍至6.16倍。
英文摘要:
Large language model (LLM) inference increasingly spans models that differ in size, capability, price, and provider. This shift creates two costs for developers. One is the integration cost of choosing among and switching between many models. The other is the inference cost of rebuilding a KV cache when it is unavailable or incompatible with the selected model. We define and analyze the Unified AI Gateway as a system setting for an edge-deployed AI traffic hub. It coordinates model routing, KV cache management, and compute placement across end devices, edge resources, and cloud model services. At request time, the gateway jointly selects a target model, an execution site, and a KV cache action under task-quality, latency, cost, and resource constraints. In parallel, background cache-management actions optimize KV cache placement, replication, retrieval, and lifecycle decisions for subsequent requests. We synthesize existing evidence on KV cache reuse, compression, cross-model mapping, distributed storage, and transfer, and discuss the remaining challenges of integrating these capabilities into one system. Across eight typical workload profiles, our workload-level analytical simulation reports TTFT speedups of 1.25$\times$--13.28$\times$ and input-cost benefits of 1.20$\times$--6.16$\times$.