arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用eBPF的微服务引导迁移的无插装依赖发现

Zero-Instrumentation Dependency Discovery for Guided Microservice Migration Using eBPF

Eshan Trivedi, Chandrahasa Pranava

arXiv 2608.04413首次发表:更新:

AI 中文总结

该研究使用eBPF内核级网络追踪无插装发现微服务依赖,生成ROI排序的迁移计划,可减少27%跨VM流量暴露,但采集开销存在明显性能代价。

AI 中文摘要

在虚拟机(VM)之间迁移微服务时,若不了解其运行时通信模式,会产生仅通过静态分析难以预测的跨VM热点和延迟峰值。我们使用扩展伯克利数据包过滤器(eBPF)内核级网络追踪,在无应用插装的情况下自动发现运行时的服务间依赖关系,并利用生成的依赖图生成按投资回报率(ROI)排序的感知流量的迁移计划。一种两次遍历的进程标识符(PID)与端口关联算法,在进程原本无法区分的共享运行时测试平台中恢复了全部20个服务的身份,匹配已知的真实拓扑。该系统从3分钟内捕获的13615个网络事件中发现了32条依赖边,并应用谱图聚类结合Kernighan-Lin优化,将服务划分为VM一致的组。在对发现的图进行的模拟中,我们按ROI排序的迁移顺序,与按字母顺序排列的、作为任意依赖盲排序的确定性代理相比,在迁移窗口内减少了27%的累计跨VM流量暴露。采集开销情况不一:在一个拥有两个虚拟CPU(vCPU)的主机上,接近饱和负载的受控A/B测试中,吞吐量仅下降4.4%,但中位数(p50)延迟上升了383%,第99百分位(p99)延迟上升了1050%。因此,我们建议在非高峰时段或专用采样节点上运行采集,而非在生产饱和负载下进行。所有结果均来自我们自行构建的单个20服务测试平台;我们不保证其在生产依赖图上的表现。

英文摘要

Migrating microservices across virtual machines (VMs) without knowledge of their runtime communication patterns risks creating cross-VM hotspots and latency spikes that are difficult to predict from static analysis alone. We use extended Berkeley Packet Filter (eBPF) kernel-level network tracing to automatically discover inter-service dependencies at runtime, with no application instrumentation, and use the resulting dependency graph to produce a traffic-aware migration plan ranked by return on investment (ROI). A two-pass process-identifier (PID) to port correlation algorithm recovers the identity of all 20 services in a shared-runtime testbed where processes are otherwise indistinguishable, matching the known ground-truth topology. The system discovers 32 dependency edges from 13,615 network events captured in three minutes, and applies spectral graph clustering with Kernighan-Lin refinement to partition services into VM-coherent groups. In simulation over the discovered graph, our ROI-ranked migration order reduces cumulative cross-VM traffic exposure during the migration window by 27% relative to alphabetical ordering, a deterministic proxy for arbitrary dependency-blind ordering. Collection overhead is mixed: in a controlled A/B test at near-saturation load on a host with two virtual CPUs (vCPUs), throughput fell by only 4.4%, but median (p50) latency rose by 383% and 99th-percentile (p99) latency by 1,050%. We therefore recommend running captures off-peak or on dedicated sampling nodes rather than under production saturation. All results are from a single 20-service testbed that we authored; we make no claim about behavior on production dependency graphs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑