arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22081cs.LGcs.AI

一种用于多类分类的无泄漏堆叠集成方法

A Leakage-Free Stacked Ensemble Method for Multiclass Classification

S. P. Sharmila, Aruna Tiwari

首次发表
浏览论文内容

中文总结 AI 辅助

针对多类分类难题,提出无泄漏堆叠集成框架LFS - FRAME,集成KAN函数学习与XGBoost规则学习,用严格策略防性能泄漏,通过学习异构基学习器概率输出提升性能,实验显示其相比单模型基线有显著准确率提升。

中文摘要 AI 辅助

多类分类是广泛领域中的一个基本问题。由于类间相似度高、数据集类不平衡以及数据分布的变异性,它仍然具有挑战性。基于规则的分类器(如XGBoost)在结构化特征上通常表现更强,但在捕捉变量间平滑函数关系方面有限。神经网络模型能表示复杂非线性交互,但常存在过拟合和泛化问题。为解决这些限制,我们提出LFS - FRAME,一种无泄漏堆叠集成框架,它集成了使用柯尔莫哥洛夫 - 阿诺德网络(KAN)的函数学习和通过XGBoost的基于规则的学习进行稳健的多类分类。该框架通过采用严格的折外堆叠策略构建无偏元特征,防止性能泄漏。通过学习异构基学习器的概率输出,元分类器有效利用了复杂数据中的全局函数模式和清晰决策边界。在多类数据集上的实验评估表明,LFS - FRAME相对于强大的单模型基线提高了性能指标,识别主要家族的总体准确率为89.85%,识别子家族的为81.74%。这些结果突出了无泄漏函数和基于规则的堆叠对可靠且可泛化的多类分类的有效性。

英文摘要

Multiclass classification is a fundamental problem across a wide range of domains. It is still challenging due to possession of high inter-class similarity, class imbalance datasets, and variability in data distributions. Rule-based classifiers such as XGBoost often achieve stronger performance on structured features, but they are limited in capturing smooth functional relationships among variables. Similarly, neural network models can represent complex nonlinear interactions but frequently suffer from overfitting and generalization issues. To address these limitations, we propose LFS-FRAME, a Leakage-Free Stacked ensemble framework that integrates functional learning using Kolmogorov-Arnold Networks (KAN) and rule-based learning via XGBoost for robust multiclass classification. The proposed framework constructs unbiased meta-features by employing a strict out-of-fold stacking strategy to ensure complete isolation between training and validation data hence preventing performance leakage. By learning over probabilistic outputs from heterogeneous base learners, the meta-classifier effectively exploits both global functional patterns and sharp decision boundaries present in the complex data. Experimental evaluations on multi-class datasets demonstrate that LFS-FRAME improves performance metrics, and overall accuracy is 89.85% in identifying major families and 81.74% in identifying sub-families relative to strong single-model baselines. These results highlight the effectiveness of leakage-free functional and rule-based stacking for reliable and generalizable multiclass classification.

发表机构

  • Indian Institute of Technology Indore(印度理工学院印多尔分校)
  • Siddaganga Institute of Technology(西达甘加理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑