arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19412cs.DC

虚假的惰性容器:Kubernetes上模型服务的惰性容器镜像拉取的延迟成本与故障语义

The Lazy Pod That Lies: Deferred Cost and Failure Semantics of Lazy Container Image Pulling for Model Serving on Kubernetes

Georgii Kliukovkin

中文总结 AI 辅助

该研究针对Kubernetes模型服务的惰性镜像拉取,发现其冷启动时间与镜像大小无关,但存在延迟成本及缓存耗尽导致的Pod故障等问题,并给出部署、监控等指导方案。

中文摘要 AI 辅助

惰性容器镜像拉取承诺通过立即挂载镜像并按需获取内容,消除启动模型服务Pod的主要成本。我们在Kubernetes上针对模型交付评估该承诺,使用KServe及两个生产级惰性拉取系统——eStargz/stargz-snapshotter和AWS SOCI,与 eager(急切)基准对比,测试对象为2至140 GB的制品,包括真实fp16权重。惰性拉取实现了其核心优势:冷启动至首次预测的时间与镜像大小无关(16.9-17.6秒,而eager拉取为24.5-573.0秒)。但成本是延迟产生而非消除:通过惰性挂载完整读取14 GB模型需105.3秒,比其替代的72.4秒eager拉取更慢,且两个系统在生命周期两端产生成本(SOCI在Pod就绪前预取几乎完整镜像;eStargz将几乎所有内容延迟至首次读取)。更关键的是,我们发现了eager拉取结构上无法出现的故障模式:在默认配置下持续合法读取时,快照器的节点级缓存耗尽其有限存储空间,已运行的Pod开始无法读取模型文件。在缓存耗尽的早期阶段,一个受检测的服务Pod通过了所有Kubernetes可见及应用层检查,持续196秒,而其快照器已记录实际故障;在更高压力下,67%-94%的模型文件读取失败,失败比例随剩余缓存占用率单调增长。若Pod下释放缓存空间,运行中的Pod会自我修复,但在运行中的Pod下重启快照器守护进程,会在仍报告为Running的Pod中留下永久陈旧的文件句柄。我们为服务平台及运维人员推导了部署、监控和缓存规模调整的指导方案。

英文摘要

Lazy container-image pulling promises to eliminate the dominant cost of starting a model-serving pod by mounting the image immediately and fetching content on demand. We evaluate this promise for model delivery on Kubernetes, using KServe with two production lazy-pulling systems -- eStargz/stargz-snapshotter and AWS SOCI -- against eager baselines, on artifacts from 2 to 140 GB including real fp16 weights. Lazy pulling delivers its headline: cold time-to-first-prediction becomes size-independent (16.9--17.6s, versus 24.5--573.0s eager). But the cost is deferred, not eliminated: a full read of a 14 GB model through the lazy mount takes 105.3s, slower than the 72.4s eager pull it replaced, and the two systems pay at opposite lifecycle ends (SOCI prefetches nearly the full image before Ready; eStargz defers nearly everything to first read). More consequentially, we characterize a failure mode eager pulling structurally cannot exhibit: under sustained legitimate reads with default configuration, the snapshotter's node-level cache exhausts its finite volume and already-running pods begin failing reads of model files. At the earliest stage of exhaustion, an instrumented serving pod passed every Kubernetes-visible and application-level check for 196s while its snapshotter was already logging real failures; under heavier pressure, 67--94% of model files fail, scaling monotonically with residual cache occupancy. A live pod self-heals if cache space is freed under it, but a snapshotter-daemon restart under a live pod leaves permanently stale file handles in a pod still reported Running. We derive placement, monitoring, and cache-sizing guidance for serving platforms and operators.

补充信息

↑