arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过同分布训练-测试划分方法增强自动机器学习

Enhancing Automated Machine Learning via Homogeneous Train-Test Splitting Methods

Yearn Tan Yin Tze, Charles Grellois

arXiv 2607.26625首次发表:更新:

发表机构

School of Computer Science, University of Sheffield(谢菲尔德大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对AutoML模型评估中训练-测试划分的分布假设偏差问题,对比多种划分策略,提出Optimised-Distribution方法,实现最高平均MMD相似性,提升模型评估准确性。

AI 中文摘要

机器学习中准确的模型评估高度依赖数据集划分为训练集和测试集的方式。标准随机划分假设两个划分具有相同的底层分布,而类别不平衡、自然聚类或空间自相关的数据集常违反该假设。本文研究训练-测试划分中的统计相似性及其对AutoML模型评估的影响,在15个UCI基准数据集上比较5种成熟策略:随机划分、分层抽样、Kennard-Stone、Duplex和SPXY,使用卡方检验、柯尔莫哥洛夫-斯米尔诺夫检验及最大均值差异(MMD)评估相似性。基于几何的方法始终产生接近零的MMD分数,导致下游性能估计不稳定。本文提出的Optimised-Distribution方法将相似性作为显式优化目标,在所有评估策略中实现最高的平均MMD相似性,达89.0%。

英文摘要

Accurate model evaluation in machine learning depends critically on how datasets are split into training and testing subsets. Standard random splitting assumes that both partitions share the same underlying distribution, an assumption often violated in datasets with class imbalance, natural clustering, or spatial autocorrelation. This paper investigates the role of statistical similarity in train-test splitting and its consequences for AutoML model evaluation. Five established strategies are compared across fifteen UCI benchmark datasets: random splitting, stratified sampling, Kennard-Stone, Duplex, and SPXY. Similarity is assessed using chi-square, Kolmogorov-Smirnov, and Maximum Mean Discrepancy (MMD) tests. Geometry-based methods consistently produce near-zero MMD scores, introducing instability into downstream performance estimates. The proposed Optimised-Distribution method treats similarity as an explicit optimisation objective and achieves the highest mean MMD similarity, 89.0%, across all strategies evaluated.

CommentsPre-review version of a paper accepted at UKCI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑