AI 中文总结
本文证明全局陈旧度阈值等价于统一定时器,提出按分段差异化刷新,最优频率与风险平方根成正比,模拟中降低陈旧暴露8-29%。
AI 中文摘要
生产环境中的机器学习模型是时间受限训练快照的派生产物:部署的模型是训练切片上的物化视图,该切片在构建的瞬间即开始老化。常见的应对方式是用自适应触发器取代固定的重训练节奏——即一种加权的陈旧度评分,当累积源风险超过阈值时触发重训练。我们证明这是错误的杠杆,并找到了正确的杠杆。首先,一个等价极限:任何作为单一共享全局训练数据年龄的静态严格单调函数的刷新触发器,在操作上等价于一个校准的统一年龄定时器,因此,无论全局陈旧度预算如何精细地对分段、来源和敏感性进行加权,它都不携带时钟所不具备的调度信息。该极限还展示了如何摆脱它:对分段进行差异化刷新,为每个分段赋予其自身的年龄和刷新间隔,当刷新成本在分段间可分离(增量训练或分段模型)时,这具有实际意义。我们解决了由此产生的预算分配问题。在频繁刷新机制下,每个分段的最优刷新率与其风险 $w_j \lambda_j$(权重乘以变化率)的平方根成正比,且最优策略的成本永远不会超过统一定时器,通过一个闭式柯西-施瓦茨“统一性代价”超越后者,该代价在同质工作负载下为零,并随异质性增长而增加。在具有真实泊松变化事件的离散事件模拟中,在匹配刷新预算下,最优策略相对于统一定时器将实现的加权陈旧暴露降低了8-29%,在86-100%的种子中获胜;而朴素的暴露阈值策略则不然,表明分配才是关键;且该优势在50%的速率估计噪声下依然存在。模型刷新的杠杆不在于更好的评分,而在于更好的行动。
英文摘要
Production machine-learning models are derived artifacts of time-bounded training snapshots: a deployed model is a materialized view over a training cut that ages the instant it is built. A common response is to replace the fixed retraining cadence with an adaptive trigger -- a weighted staleness score that retrains when accumulated source risk crosses a threshold. We show this is the wrong lever, and identify the right one. First, an equivalence limit: any refresh trigger that is a static, strictly monotone function of a single shared global training-data age is operationally equivalent to a calibrated uniform age timer, so a global staleness budget, however elaborately it weights segments, sources, and sensitivities, carries no scheduling information a clock does not. The limit also shows how to escape it: refresh segments differentially, giving each its own age and refresh interval, which is meaningful when refresh cost is separable across segments (incremental training or per-segment models). We solve the resulting budget-allocation problem. In the frequent-refresh regime each segment's optimal refresh rate is proportional to the square root of its risk $w_j λ_j$ (weight times change rate), and the optimal policy never costs more than the uniform timer, beating it by a closed-form Cauchy-Schwarz "price of uniformity" that is zero for homogeneous workloads and grows with heterogeneity. In a discrete-event simulation with real Poisson change events, the optimal policy lowers realized weighted stale exposure by 8-29% relative to the uniform timer at matched refresh budget, winning on 86-100% of seeds; a naive exposure-threshold policy does not, showing the allocation is what helps; and the advantage survives 50% rate-estimation noise. The leverage in model refresh is not a better score but a better action.