arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09432cs.LGcs.AI

SCCM:用于自动漂移检测与适应的流巡航控制方法

SCCM : Stream Cruise Control Method for Automated Drift Detection and Adaptation

  • University of North Texas(北德克萨斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Mohammad Abu-Shaira, Weishi Shi

AI总结:

SCCM提出一种流巡航控制方法,通过早期漂移检测、动态阈值和模型重校准,实现在线回归中的自动漂移适应,在合成与真实数据上优于基线。

AI中文摘要:

真实世界的数据集通常表现出不断演变的分布,即所谓的概念漂移。忽略漂移会降低预测性能,而对固定超参数的依赖进一步限制了模型在变化条件下的适应性。自适应学习通过在线持续更新模型来应对这一挑战,使模型能够随着数据分布的演变逐步调整并保持有效性。本文提出了流巡航控制方法(SCCM),一个用于在线回归中漂移检测与适应的综合框架。SCCM通过早期响应、更新前漂移检测、漂移幅度量化、基于KPI窗口的阈值设定以缓解局部误报、动态超参数调整以及模型重新校准,实现了自动化适应。与通常在观察到性能下降后才激活适应的纯反应式方法不同,SCCM采用内存内设计以实现实时适应性。通过使用动态阈值并保持对数据分布的中立性,SCCM支持跨不同数据流(包括高维和大规模场景)的基于KPI的监控。SCCM与四种在线回归模型集成,并在覆盖突变、增量及交替渐变漂移的18个合成数据集以及八个真实世界数据集上进行了评估。评估同时使用R2和MSE指标,并与八个检测器-适应基线进行了比较。结果表明,在所评估的在线回归设置中,SCCM提升了预测性能并有效处理了漂移。

英文摘要:

Real-world datasets often exhibit evolving distributions, known as concept drift. Ignoring drift degrades predictive performance, while reliance on fixed hyperparameters further limits model adaptability under changing conditions. Adaptive learning addresses this challenge by continuously updating models online, allowing them to incrementally adjust and remain effective as data distributions evolve. This paper presents the Stream Cruise Control Method (SCCM), a comprehensive framework for drift detection and adaptation in online regression. SCCM enables automated adaptation through early-response, pre-update drift detection, drift magnitude quantification, KPI-window-based thresholding for local false-alarm mitigation, dynamic hyperparameter tuning, and model recalibration. SCCM also adopts an in-memory design for real-time adaptability, unlike purely reactive methods that typically activate adaptation only after performance degradation is observed. By using dynamic thresholding and remaining agnostic to data distributions, SCCM supports KPI-based monitoring across varying data streams, including high-dimensional and large-scale settings. SCCM is integrated with four online regression models and evaluated on 18 synthetic datasets covering abrupt, incremental, and alternating gradual drift, together with eight real-world datasets. The evaluation uses both R2 and MSE and compares against eight detector--adaptation baselines. Results show improved predictive performance and effective drift handling across the evaluated online regression settings.

↑