在共置的轻量级虚拟机上提供微秒级跨虚拟机核心弹性
Offering Microsecond-Scale Cross-VM Core Elasticity on Colocated Lightweight Virtual Machines
浏览论文内容
中文总结 AI 辅助
该研究提出基于 KVM 的超轻量级虚拟机 substrate HyperFlux,实现微秒级跨虚拟机核心弹性,可降低高优先级虚拟机尾部延迟,优于传统方案。
中文摘要 AI 辅助
无服务器平台通常会将多种不同工作负载共置在快速启动、内存占用小的虚拟机(VM)中,以提高部署密度。为应对流量突发时的尾部延迟,为每个虚拟机配置超出其需求的资源会损害密度;而要在保持高密度的同时有效保护尾部延迟,基础设施需能以微秒级时间尺度将物理核心分配给任何受延迟敏感的突发虚拟机,并在突发结束后收回核心。目前尚无虚拟机 substrate 能实现这一点:传统虚拟机仅能通过毫秒级的 vCPU 热插拔路径调整客户机核心数量,Firecracker 在启动时就固定了虚拟机的核心数,而启动最快的超轻量级虚拟机则完全放弃了多核执行。我们提出 HyperFlux,这是一种基于商用 KVM 的超轻量级虚拟机 substrate,可在运行时使虚拟机的并行度(支撑它的物理核心数)具备弹性。研究表明,HyperFlux 仅需 13μs 即可在虚拟机间迁移一个核心,即使从繁忙的捐赠虚拟机中强制收回核心,速度也比 vCPU 热插拔快几个数量级。HyperFlux 虚拟机仅占用 3.2MB 内存,冷启动时间为 1.37ms,与启动最快的超轻量级虚拟机相当,且独特支持多核并行。在共置场景下,与 Firecracker 和 Cloud Hypervisor 的静态核心共享相比,它在高负载下可将高优先级虚拟机的尾部延迟降低多达 10 倍,且在负载突发变化时,与使用 cgroup 和 vCPU 热插拔相比,能提供更低、更稳定的尾部延迟。
英文摘要
Serverless platforms commonly colocate many diverse workloads, each in a fast-booting, memory-lean virtual machine (VM), to improve deployment density. Overprovisioning each VM for its peak protects tail latency during traffic bursts but hurts density; maintaining high density while effectively protecting tail latency requires the infrastructure to be able to shift physical cores, at a microsecond timescale, to whichever latency-sensitive VM is bursting and reclaim them as the burst subsides. No VM substrate delivers this: conventional VMs resize a guest's cores only through a millisecond-scale vCPU hot-plug path, Firecracker fixes a VM's core count at boot, and the ultralight VMs that boot fastest drop multicore execution entirely. We present HyperFlux, a commodity-KVM ultralight VM substrate that makes a VM's parallelism width (the number of physical cores backing it) elastic at runtime. We show that HyperFlux can move a core across VMs in merely 13$μ$s, even when forcibly reclaiming it from a busy donor, orders of magnitude faster than vCPU hot-plug. A HyperFlux VM incurs only a 3.2MB memory footprint and can cold-boot in 1.37ms, on par with the fastest-booting ultralight VMs, while uniquely supporting multicore parallelism. Under colocation, it can reduce high-priority VMs' tail latency by up to 10x under high load compared to static core-sharing with Firecracker and Cloud Hypervisor, and deliver a lower and more stable tail latency compared to using cgroup and vCPU hot-plug under changing load bursts.