AI 中文总结
该研究对比三种模型刷新策略与不重新训练基准,发现增量学习是关键,无增量时定期重新训练在多数场景优于反应式策略,还揭示反应式策略失效模式及延迟预算交互问题。
AI 中文摘要
生产型机器学习系统会因概念漂移而性能下降,但从业者对于何时重新训练几乎没有原则性指导。重新训练成本高昂,重新训练预算有限,且重新训练后的模型无法立即生效:训练与部署延迟会导致陈旧模型在数据持续变化的情况下提供预测。我们针对三种实用的模型刷新策略(定期重新训练、误差阈值触发、结合ADWIN的统计漂移触发重新训练)与不重新训练的基准策略开展受控实证研究,采用统一系统模型明确重新训练预算及训练加部署延迟。在涵盖三种漂移场景、三个预算水平、最高五个延迟水平、三个数据集及两种学习模式的3933次实验运行中,我们发现最关键的设计决策并非重新训练策略,而是部署的模型是否进行增量学习。对于每样本增量更新,以及本研究中采用即时标签的线性在线学习器而言,在54组配对比较中,即使处于极端延迟下,任何策略与不重新训练基准之间也无实际显著差异。若无增量更新,策略选择会使漂移后准确率产生15至55个百分点的结果差异,且简单的定期重新训练在突发漂移与渐进漂移场景下显著优于两种反应式策略,而反应式策略仅在重复漂移场景下保留优势。我们记录了反应式策略的系统性失效模式,以及延迟-预算排队交互会使有效重新训练预算减半的现象,并发布了完整模拟器、数据集流水线及每次运行的人工制品以确保可复现性。
英文摘要
Production machine learning systems degrade under concept drift, yet practitioners have little principled guidance on when to retrain. Retraining is costly, retraining budgets are finite, and a retrained model does not take effect instantly: training and deployment latency leave a stale model serving predictions while the data continues to move. We present a controlled empirical study of three practical model-refresh policies (periodic retraining, error-threshold triggering, and statistical drift-triggered retraining with ADWIN) against a no-retrain baseline, evaluated under a unified system model that makes retraining budgets and training-plus-deployment latency explicit. Across 3,933 experiment runs spanning three drift regimes, three budget levels, up to five latency levels, three datasets, and two learning modes, we find that the single most consequential design decision is not the retraining policy but whether the deployed model learns incrementally. With per-sample incremental updates, and for the linear online learner with immediate labels studied here, no policy differs from the no-retrain baseline by a practically significant margin in any of 54 paired comparisons, even at extreme latency. Without incremental updates, policy choice separates outcomes by 15-55 percentage points of post-drift accuracy, and simple periodic retraining significantly outperforms both reactive policies under abrupt and gradual drift, while reactive policies retain an advantage only under recurring drift. We document systematic failure modes of reactive policies and a latency-budget queueing interaction that silently halves effective retraining budgets, and release the full simulator, dataset pipelines, and per-run artifacts for reproducibility.
Comments16 pages, 7 figures, 7 tables. Code and data: https://github.com/sa1dasari/Study-of-Drift-Triggered-Retraining-Policies-Under-Budget-and-Latency-Constraints