AI 中文总结
研究在机器学习服务中,不可信学习者可能导致服务水平协议失效的问题。提出用小型可信防护机制包裹不可信学习者,将租户服务水平协议义务分两部分处理,通过实验验证该机制能有效保证服务质量,还刻画了廉价静态筛选的可信情况。
AI 中文摘要
现代机器学习服务越来越多地让经过学习的、无界的组件(路由器、延迟服务水平协议准入器、准入阶梯)来决定租户的服务质量;一旦出现错误,保证的服务水平协议可能会悄然失效,而其下方的Kubernetes层(Kueue、DRA、网关API推理扩展)只会带来跨层的意外情况。我们不是信任学习者正确,而是限制错误学习者可能造成的损害:一个小型可信防护机制包裹不可信学习者——学习者提出建议,防护机制进行处理。租户的保证服务水平协议义务分为具有不同认知的两部分。其安全预测(每个请求的调度可行性以及每个类、每个窗口的服务下限)是防护机制在运行时强制执行的可控义务,通过预留和优先级调度,无论学习者准入器出现多么严重的错误。其总体义务(尾部延迟百分位数)是一个统计残差,没有每个请求的执行点;针对它的廉价确定性筛选不仅在接近饱和时,而且在远低于饱和时都是乐观不合理的。在真实的2xV100上,在两种方向相反的舍弃策略下,对于学习者准入器的每一次错误校准,防护机制的保证类错误率为0.0(重复10次,最坏情况下限威尔逊置信区间为0.0053)。在服务模拟器上针对实际部署的GAIE流量控制,一个错误标记的路由器将相同的保证请求的错误率从0.0变为1.0;我们的防护机制按真实类别进行预留,因此标签不会打破其下限。我们描述了何时可以信任廉价的静态筛选以及何时不能。作为一篇前沿风格的论文,我们在商品2xV100和服务模拟器上评估了相关立场,并将数据中心规模、真实模型流量控制和一个封闭的最坏情况定理作为议程。
英文摘要
Modern ML serving increasingly lets learned, unverified components (routers, latency-SLO admitters, admit ladders) decide a tenant's quality of service; when one is wrong, the assured SLO can silently break, and the Kubernetes layers beneath (Kueue, DRA, the Gateway-API Inference Extension, GAIE) add cross-layer surprises. Rather than trust the learner to be right, we bound the damage a wrong one can do: a small trusted guard wraps the untrusted learner (learned proposes, the guard disposes). A tenant's assured-SLO obligation splits into two parts with different epistemics. Its safety projection, a per-class, per-window assured floor (with an optional drop rule, doom-sound only under an assumed service lower envelope), is a controllable obligation a guard enforces at runtime, holding it regardless of a learned admitter that is arbitrarily wrong within a bounded proposal interface. The admission floor is enforced structurally; given the stated assumptions, the service floor follows as a conditional response-time implication. Its aggregate obligation (the population tail-latency percentile) has no per-request enforcement point, so we treat it as a statistical residual and screen it. On real 2xV100 the guard (a Simplex-style assured-floor gate plus assured-first priority) holds assured-class miss 0.0 across every tested miscalibration of a learned admitter that, unguarded, misses 0.86-0.94; against a live deployment of the GAIE Flow Control, an injected mapping fault (emulating an untrusted mapper) flips the same assured requests from miss 0.0 to 1.0 (a mechanism-level trust-boundary test, not a head-to-head), while our guard reserves by the true class. As a Frontiers submission we evaluate the stance on commodity 2xV100 and a serving simulator, scoping datacenter scale, real-model Flow Control, and a closed worst-case theorem as the agenda.
Comments12 pages, 3 figures