Kubernetes推理服务中的冷启动模型交付:基于OCI的分发及其完整性的实证研究
Cold-Start Model Delivery in Kubernetes Inference Serving: An Empirical Study of OCI-Based Distribution and Its Integrity
浏览论文内容
中文总结 AI 辅助
研究Kubernetes推理服务中模型冷启动交付问题,通过在KServe平台实现oci+native://和oci+fetch://两条新交付路径进行分析,对比不同交付路径在不同模型大小下的表现,提出服务时完整性设计并验证其对交付时间的影响。
中文摘要 AI 辅助
Kubernetes上模型服务Pod的启动延迟主要由模型权重交付这一步骤决定。随着模型权重达到大语言模型的数百GB,冷启动交付时间决定了自动缩放和缩至零的经济性,但主要机制仍是从对象存储进行临时下载,没有利用Kubernetes为容器镜像提供的拉取缓存、摘要寻址或验证功能。我们从两个轴分析Kubernetes服务平台可用的交付路径:哪个组件拉取工件,以及是否有准入时验证器可以将部署的引用绑定到到达的字节。我们通过在广泛部署的CNCF模型服务平台KServe中实现两条新的交付路径来验证上游分析:oci+native://,它将模型镜像挂载为Kubernetes镜像卷(KEP-4639,已合并到上游),以及oci+fetch://,它在存储初始化器中拉取OCI工件(正在审核)。我们报告了据我们所知在Kubernetes服务平台上对模型交付路径(模型车边车、原生镜像卷、对象存储下载)进行的首次受控比较,工件大小为1B、7B和70B类模型的fp16权重(2-140GB)。节点缓存的OCI交付使热副本添加与大小无关:70B类工件为11.7秒,而通过对象存储重新下载则需要40.7分钟,相差208倍,而首次冷拉成本最高可达普通下载的2倍,局限于容器运行时的blob写入然后解包双程。对于s3://、gs://或hf:// URI上的模型,在没有准入时验证器观察字节的情况下,我们向KServe社区提出了一种服务时完整性设计:在存储初始化器中进行摘要固定和OpenSSF模型签名强制。下载期间的流哈希验证使交付时间增加不到0.1%;下载后检查最多增加53%。
英文摘要
The startup latency of a model-serving pod on Kubernetes is dominated by one step: delivering the model weights. As models reach the hundred-gigabyte weights of large language models, cold-start delivery time governs the economics of autoscaling and scale-to-zero, yet the dominant mechanisms remain ad-hoc downloads from object storage, with none of the pull caching, digest addressing, or verification Kubernetes provides for container images. We analyze the delivery paths available to a Kubernetes serving platform along two axes: which component pulls the artifact, and whether any admission-time verifier can bind the deployed reference to the arriving bytes. We validate the analysis upstream in KServe, a widely deployed CNCF model-serving platform, by implementing two new delivery paths: oci+native://, which mounts model images as Kubernetes image volumes (KEP-4639), merged upstream, and oci+fetch://, which pulls OCI artifacts inside the storage initializer, under review. We report, to our knowledge, the first controlled comparison of model delivery paths in a Kubernetes serving platform (modelcar sidecars, native image volumes, object-storage download) on artifacts sized to fp16 weights of 1B-, 7B-, and 70B-class models (2-140 GB). Node-cached OCI delivery makes warm replica addition size-independent: 11.7 s for a 70B-class artifact versus 40.7 minutes of re-download over object storage, a 208x difference, while the first cold pull costs up to 2x a plain download, localized to containerd's blob-write-then-unpack double pass. For models on s3://, gs://, or hf:// URIs, where no admission-time verifier observes the bytes, we present a serving-time integrity design proposed to the KServe community: digest pinning and OpenSSF model-signing enforcement in the storage initializer. Streaming hash verification during download adds under 0.1% to delivery time; a post-download pass adds up to 53%.