arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38235cs.NI

基于 eBPF 的多集群 Kubernetes 数据平面特征化:Cilium Cluster Mesh 与 KVStoreMesh 的系统性评估

Characterizing the eBPF-Based Data Plane for Multi-ClusterKubernetes: A Systematic Evaluation of Cilium Cluster Meshand KVStoreMesh

Simhadri Podala Narasimha

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对多集群 Kubernetes 中基于 eBPF 的数据平面(Cilium Cluster Mesh)缺乏系统评估的问题,提出包含五个研究问题的方法论,并给出校准模拟的初步预测结果,为实证研究提供模板。

中文摘要 AI 辅助

Kubernetes 部署日益跨越多个集群,原因包括规模、故障隔离、监管边界以及异构硬件部署,包括专用于 AI 和 HPC 工作负载的 GPU 密集型集群。Cilium Cluster Mesh 构建在扩展的 Berkeley 数据包过滤器(eBPF)技术之上,已成为一种广泛部署的机制,用于将此类集群连接成单一逻辑网络,而无需专用的多集群网关。尽管其被采用,学术文献中尚无对基于 eBPF 的多集群数据平面本身的系统性、可复现评估:现有工作要么在单个集群内将 eBPF 与 iptables 进行基准测试,要么将 eBPF 应用于跨集群监控而非转发和身份传播。可用的最详细规模数据——一份关于 clustermesh-apiserver 在约 45,000 个节点、256 个集群的负载下发生故障的报告——来自行业工程博客而非同行评审来源。本文提出了一种方法论来填补这一空白。我们定义了五个研究问题,涵盖延迟、吞吐量、控制平面限制、扰动以及 GPU/RDMA 敏感性,并描述了测试平台和 eBPF 级仪器设计。由于本次投稿缺乏实时多集群测试平台,我们报告了针对生产数据点校准的控制平面架构的校准分析模拟结果。这些结果以可证伪的预测形式呈现,旨在测试平台完成后进行验证。本文提供了动机、实验设计和初步模拟结果,作为完整实证研究的模板。

英文摘要

Kubernetes deployments increasingly span multiple clusters for reasons of scale, fault isolation, regulatory boundaries, and heterogeneous hardware placement, including GPU-dense clusters dedicated to AI and HPC workloads. Cilium Cluster Mesh, built on extended Berkeley Packet Filter (eBPF) technology, has emerged as a widely deployed mechanism for connecting such clusters into a single logical network without a dedicated multi-cluster gateway. Despite its adoption, the academic literature contains no systematic, reproducible evaluation of the eBPF-based multi-cluster data plane itself: existing work either benchmarks eBPF against iptables within a single cluster, or applies eBPF to cross-cluster monitoring rather than forwarding and identity propagation. The most detailed scale data available - a report of a clustermesh-apiserver failing under load at approximately 45,000 nodes across 256 clusters - comes from an industry engineering blog rather than a peer-reviewed source. This paper proposes a methodology to close that gap. We define five research questions covering latency, throughput, control-plane limits, churn, and GPU/RDMA sensitivity, and describe a testbed and eBPF-level instrumentation design. Lacking a live multi-cluster testbed for this submission, we report results from a calibrated analytical simulation of control-plane architectures fit to the production data point. These results are presented as falsifiable predictions intended for validation once the testbed is complete. This paper provides the motivation, experimental design, and preliminary simulation results to serve as a template for a full empirical study.

补充信息

↑