arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

因果视角下的概念漂移

Concept Drift from a Causal Perspective

Eduardo V. L. Barboza, Jean Paul Barddal, Robert Sabourin, Rafael M. O. Cruz

arXiv 2609.25340首次发表:更新:

发表机构

LIVIA, École de Technologie Supérieure; Graduate Program in Informatics (PPGIa), Pontifícia Universidade Católica do Paraná (PUCPR)(LIVIA,高等工程技术学院; 信息学研究生项目(PPGIa),巴拉那天主教大学(PUCPR))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文从因果视角重新定义概念漂移,提出基于结构因果模型的分类法及数据流生成器,实验表明不同因果来源的漂移产生不同分布效应,并验证了因果感知评估的价值。

AI 中文摘要

概念漂移是现实世界数据流中的常见现象,其中数据生成分布的变化会降低预测模型的性能。大多数现有定义将漂移描述为联合分布$P(\mathbf{x}, y)$的变化,而未区分数据生成过程中哪个组成部分发生了变化。在本工作中,我们基于结构因果模型(SCMs)引入了概念漂移的因果视角。我们提出了一种分类法,根据因果来源对漂移事件进行分类,包括外生变量、内生机制、混杂因素和目标生成过程的变化。在此框架基础上,我们开发了一个基于SCM的数据流生成器,用于模拟受控的机制级漂移事件。我们的实验从经验上刻画了每种漂移类型的分布效应,并表明不同因果来源的漂移会引发不同的分布偏移模式和预测行为。此外,通过整合因果发现方法,我们利用该框架构建了基于真实世界依赖结构的数据流,从而实现了更真实且信息量更大的评估场景。我们还证明了利用生成的数据可以提升下游性能。这些结果凸显了在研究评估自适应学习方法时考虑因果结构的重要性,并为非平稳环境中的因果感知评估奠定了基础。

英文摘要

Concept drift is a common phenomenon in real-world data streams, in which changes in the data-generating distribution can degrade predictive model performance. Most existing definitions characterize drift as changes in the joint distribution $P(\mathbf{x}, y)$, without distinguishing which component of the data-generating process has changed. In this work, we introduce a causal perspective on concept drift based on Structural Causal Models (SCMs). We propose a taxonomy that categorizes drift events by their causal origin, including changes in exogenous variables, endogenous mechanisms, confounders, and target-generating processes. Building on this framework, we develop an SCM-based data stream generator that simulates controlled mechanism-level drift events. Our experiments empirically characterize the distributional effects of each drift type and show that drifts with different causal origins induce distinct patterns of distribution shift and predictive behavior. Furthermore, by integrating causal discovery methods, we use our framework to construct data streams grounded in real-world dependency structures, enabling more realistic and informative evaluation scenarios. We also demonstrate that leveraging the generated data can improve downstream performance. These results highlight the importance of accounting for causal structure when studying and evaluating adaptive learning methods, and establish a foundation for causally-aware evaluation in non-stationary environments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑