arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16659cs.LGcs.AI

面向带概念漂移的数据流分类与集成学习的霍夫丁自适应分裂树

Hoeffding adaptive splitting trees for data stream classification with concept drift and ensemble learning

  • Pontifícia Universidade Católica do Paraná (PUCPR)(巴拉那天主教大学)
  • Sorbonne Université(索邦大学)

机构由 AI 辅助整理,请以论文原文为准。

Daniel Nowak Assis, Jean Paul Barddal, Fabrício Enembreck

AI总结:

本文针对自适应分裂决策树作为集成基学习器的多样性不足问题,提出霍夫丁自适应分裂树模型,结合霍夫丁树的周期性分裂策略与自适应分裂机制,在综合评估中实现了最优的数据流分类性能。

AI中文摘要:

决策树集成是数据流分类领域成熟的方法,在集成学习中,霍夫丁树(Hoeffding Trees)作为基学习器被广泛采用,其依据霍夫丁界进行周期性分裂尝试。然而近期研究表明,这种标准分裂机制缺乏适应性,而针对性能下降触发分裂的自适应树已取得更优结果。本文指出自适应分裂决策树作为集成基学习器存在局限性,发现变化检测器往往无法提升集成内部的充分多样性。为解决该问题,我们提出两种新型决策树模型,即霍夫丁自适应分裂树(Hoeffding Adaptive Splitting Trees)。这些模型结合了霍夫丁树用于促进集成多样性的周期性分裂策略,以及采用变化检测算法识别性能衰减并确定分裂点的自适应分裂机制。实验结果表明,霍夫丁自适应分裂树可提升集成性能,并在包含基准对比、计算成本分析及概念漂移适应的综合评估中达到了当前最优结果。

英文摘要:

Ensembles of decision trees are well-established methods for data stream classification. In ensemble learning, Hoeffding Trees are widely adopted as base learners, performing periodic split attempts according to the Hoeffding bound. Recent studies, however, indicate that this standard splitting mechanism lacks adaptability, while adaptive trees that trigger splits in response to performance degradation have achieved superior results. In this paper, we identify limitations in the use of adaptive-splitting decision trees as ensemble base learners, showing that change detectors often fail to promote sufficient diversity within ensembles. To address this issue, we propose two novel decision tree models, termed Hoeffding Adaptive Splitting Trees. These models combine the periodic splitting strategy of Hoeffding Trees, which fosters ensemble diversity, with adaptive splitting mechanisms that employ change detection algorithms to identify performance decay and determine split points. Experimental results demonstrate that Hoeffding Adaptive Splitting Trees enhance ensemble performance and achieve state-of-the-art results across a comprehensive evaluation, including benchmark comparisons, computational cost analysis, and concept drift adaptation.

↑