AI 中文总结
针对多云Kubernetes集群的API网关部署,本文提出MILP模型与贪心算法,权衡成本、资源与延迟,贪心算法加速比高且成本接近最优,但部分场景无法满足延迟要求。
AI 中文摘要
在地理分布式的多云Kubernetes集群中部署API网关时,需在基础设施成本、计算资源与网络延迟之间进行权衡。本文提出一种优化模型,将API网关部署问题转化为带容量约束的设施定位问题,联合确定需激活的候选集群、部署的网关副本数量,以及区域流量在选定集群间的分配方式。该模型对客户端到集群的网络往返延迟设定上限(不包含网关处理、排队及后端服务延迟),并纳入网关副本容量的利用率余量因子。本文同时提出混合整数线性规划(MILP)模型与构造性贪心启发式算法,该算法依据增量成本对候选集群排序,增量成本包含集群激活成本与边际副本成本,且需结合已分配负载,按可分配容量单位计算。两种模型均应用于确定性、种子控制的地理合成实例,针对每种问题规模生成30个带随机种子的实例以分析性能。贪心算法相较于MILP最优解的最优性间隙为3.2%至4.7%,某一特定实例的最大间隙达25.0%,在3至12个候选集群的场景下,加速比约为660倍至3490倍。在典型的10个候选集群、10个需求区域的实例中,MILP最优部署方案相较全副本基线方案可节省24.2%的月度成本;而选择单个最便宜的候选集群虽相较MILP最优方案节省24.8%成本,但无法满足10个需求区域中3个区域的延迟要求。
英文摘要
The use of API gateways within geographically distributed multi-cloud Kubernetes clusters poses a tradeoff between infrastructure cost, computational resources, and network latencies. We present an optimization formulation that addresses API gateway placement as a capacitated facility location problem that jointly determines which candidate clusters to activate, how many gateway replicas to deploy, and how regional traffic should be distributed across the selected clusters. The formulation imposes an upper bound on estimated client-to-cluster network round-trip latency, excluding gateway processing, queuing, and backendservice latency, and incorporates a utilization headroom factor for gateway replica capacity. We present both a mixed-integer linear programming (MILP) formulation and a constructive greedy heuristic that ranks candidates according to incremental cost, comprising cluster-activation and marginal replica costs, per unit of assignable capacity while explicitly accounting for already committed load. Both formulations are applied to deterministic, seed-controlled, geography-based synthetic instances. For each problem size, 30 instances are generated with random seeds to analyze their performance. The greedy algorithm achieves an optimality gap of 3.2% to 4.7% to the MILP optimal solution, with a maximum observed gap of 25.0% for one particular instance, and a speedup of approximately 660x to 3,490x for 3 to 12 candidate clusters. In a canonical 10-candidate, 10-demand region instance, MILP-optimal deployment saves 24.2% in terms of monthly cost compared to the full-replication baseline. On the other hand, selecting the single cheapest candidate yields savings of 24.8% compared to the MILP optimum but does not satisfy the latency requirement for 3 out of 10 demand regions.