arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

演化概念定义下的溯源引导增量学习

Provenance Guided Incremental Learning Under Evolving Concept Definitions

Ismail Lamaakal

arXiv 2608.23893首次发表:更新:

发表机构

Mohammed Premier University(穆罕默德一世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对规则诱导的概念漂移,提出溯源引导增量学习框架,结合RuleShift-Bench基准验证,该方法仅重新处理少量历史数据即可实现高精度的概念修订修复,大幅降低更新延迟。

AI 中文摘要

长期部署的学习系统不仅要适应传入数据的统计变化,还要调整生成预测目标的定义。传统概念漂移方法通常从观测或预测误差中推断此类变化,即便底层策略、规则或查询已被明确修改。本文研究规则诱导的概念漂移,即定义目标的概念被直接修改,导致先前存储的实例在观测数据未变的情况下获得不同语义标签。我们提出溯源引导增量学习框架,该框架将连续的概念定义编译为结构化规则增量,通过历史溯源追踪变化组件,认证先前标签仍有效的记录,并将重新评估限制在局部候选区域。可执行修订自动重新标注,模糊案例通过选择性监督处理,所得变更用于增量预测器修复。版本化概念记忆进一步支持重复出现的概念定义。我们还推出RuleShift-Bench,涵盖金融、人口统计、网络安全和图结构数据,包含阈值、谓词、逻辑、关系、重复及混合概念修订。在该基准上,溯源引导修复达到92.3%的准确率和90.2%的Macro-F1,仅重新处理14.7%的历史集合,保留94.6%的受影响记录,其平均更新延迟为179秒,而完全重新标注和重新训练的延迟为993秒。结果表明,明确的概念修订可被用作数据维护信号,使学习系统能更新依赖该变更的监督和预测状态,同时保留仍有效的知识。

英文摘要

Learning systems deployed over long periods must adapt not only to statistical changes in incoming data, but also to revisions of the definitions that generate their prediction targets. Conventional concept-drift methods typically infer such changes from observations or prediction errors, even when the underlying policy, rule, or query has been explicitly modified. This paper studies rule-induced concept shift, where the target-defining concept is revised directly, causing previously stored instances to acquire different semantic labels without requiring any change in their observed data. We introduce a provenance-guided incremental learning framework that compiles consecutive concept definitions into a structured rule delta, traces the changed components through historical provenance, certifies records whose previous labels remain valid, and restricts reevaluation to a localized candidate region. Executable revisions are relabeled automatically, ambiguous cases are handled through selective supervision, and the resulting changes are used for incremental predictor repair. A versioned concept memory further supports recurring definitions. We also introduce RuleShift-Bench, spanning financial, demographic, cybersecurity, and graph-structured data with threshold, predicate, logical, relational, recurring, and mixed concept revisions. Across the benchmark, provenance-guided repair attains 92.3% accuracy and 90.2% Macro-F1 while reprocessing 14.7% of the historical collection and retaining 94.6% of affected records. Its average update latency is 179s compared with 993s for complete relabeling and retraining. The results demonstrate that an explicit concept revision can be exploited as a data-maintenance signal, allowing learning systems to update the supervision and predictive state that depend on the change while preserving knowledge that remains valid.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑