arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EasyClassifier:为不编程的研究者提供的诚实、可复现的机器学习分类工具

EasyClassifier: Honest, Reproducible Machine-Learning Classification for Researchers Who Do Not Program

Ahmad B. A. Hassanat, Ghada A. Altarawneh

arXiv 2610.04758首次发表:更新:

发表机构

Mutah University(穆塔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EasyClassifier 是一个面向非编程研究者的开源 Python 包,通过默认正确的流程避免数据泄漏和选择偏差,在十个数据集上显著降低准确率高估,并提供可复现的诚实结果。

AI 中文摘要

机器学习分类现已广泛应用于医学、社会科学、经济学、教育和工程领域,使用者往往是不从事编程的研究者。此类工作中常出现两种错误:在交叉验证之前,基于所有行学习预处理(数据泄漏);以及报告在比较中胜出的分类器的交叉验证得分(选择偏差)。两者都会使发表的得分过于乐观。我们提出了 EasyClassifier,一个开源的 Python 包,通过一系列通俗易懂的多项选择题,引导用户从 CSV 或 Excel 文件完成分析,无需编写任何代码。它将正确的流程设为默认选项而非可选:每个预处理步骤都在每个训练折内学习;选定的分类器通过嵌套交叉验证(最多 2000 行)或未触碰的 20% 测试集获得独立的最终得分;分类器使用固定的默认参数和固定的随机种子运行;每次运行都会生成一份报告,包含可直接修改的方法段落、需引用的参考文献以及可发表的图表。在来自七个领域的十个公开数据集和一个随机标签对照上,将数据对半分割,一半作为未触碰的外部测试集,常见做法平均高估平衡准确率 2.9 个百分点(在小数据上最高达 9.9 个百分点),而 EasyClassifier 报告的得分平均有符号偏差为 +0.7 个百分点,平均绝对误差为 2.3 个百分点(单侧 Wilcoxon p = 3.2 x 10^{-4},30 次运行)。在随机标签上,它报告的平衡准确率为 48.8%,而随机水平为 50%。EasyClassifier 以 MIT 许可证发布,可从 PyPI(pip install easyclassifier)、GitHub 和 Zenodo 获取。

英文摘要

Machine-learning classification is now used across medicine, the social sciences, economics, education and engineering, very often by researchers who do not program. Two errors recur in such work: preprocessing learned on all rows before cross-validation (data leakage), and reporting the cross-validation score of the classifier that won a comparison (selection bias). Both make published scores optimistic. We present EasyClassifier, an open-source Python package that guides a user from a CSV or Excel file to a finished analysis through a sequence of plain-language, multiple-choice questions, with no code. It makes the correct procedure the default rather than an option: every preprocessing step is learned inside each training fold; the selected classifier receives a separate final score from nested cross-validation (up to 2,000 rows), or an untouched 20% test set; classifiers run with fixed default parameters and fixed random seeds; and every run writes a report with a ready-to-adapt Methods paragraph, the references to cite, and figures prepared for publication. On ten public datasets from seven fields and a random-label control, split in half so that one half served as an untouched external test, the common practice overstated balanced accuracy by 2.9 percentage points on average (up to 9.9 on small data), whereas the score EasyClassifier reports had a mean signed bias of +0.7 points and a mean absolute error of 2.3 points (one-sided Wilcoxon p = 3.2 x 10^{-4}, 30 runs). On random labels it reported 48.8% balanced accuracy against a chance level of 50%. EasyClassifier is available under the MIT license from PyPI (pip install easyclassifier), GitHub, and Zenodo.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑