发表机构
Gdansk University of Technology(格但斯克工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估Canny边缘检测预处理对帕金森病分类ML模型的影响,发现该预处理降低多数模型性能,但随机森林内存稳定,逻辑回归等模型表现稳健。
AI 中文摘要
本研究探讨了使用机器学习(ML)模型将个体分类为健康或帕金森病风险人群的问题,重点关注数据集规模和预处理技术对模型性能的影响。从原始数据集创建了四个数据集:DS_0(正常数据集)、DS_1(对DS_0进行Canny边缘检测和Hessian滤波处理)、DS_2(增强的DS_0)和DS_3(增强的DS_1)。我们评估了一系列ML模型——逻辑回归(LR)、决策树(DT)、随机森林(RF)、梯度提升(GB)、XGBoost(XBG)、朴素贝叶斯(NB)、支持向量机(SVM)和AdaBoost(AdB)——在这些数据集上的表现,分析预测准确率、模型大小和预测延迟。结果表明,虽然更大的数据集导致模型内存占用和预测延迟增加,但由Hessian滤波补充的Canny边缘检测预处理(用于DS_1和DS_3)降低了大多数模型的性能。在我们的实验中,观察到随机森林(RF)在所有数据集上保持稳定的61 KB内存占用,而KNN和SVM等模型的内存使用显著增加,从DS_0上的5.7-7 KB增加到DS_2上的102-220 KB,预测时间也类似增加。逻辑回归、决策树和朴素贝叶斯在所有数据集上表现出稳定的内存占用和快速的预测时间。XGBoost的预测时间从DS_0上的180-200毫秒增加到DS_2上的700-3000毫秒(截断)。
英文摘要
This study investigates the classification of individuals as healthy or at risk of Parkinson's disease using machine learning (ML) models, focusing on the impact of dataset size and preprocessing techniques on model performance. Four datasets are created from an original dataset: DS_0, (normal dataset), DS_1 (DS_O subjected to Canny edge detection and Hessian filtering), DS_2 (augmented DS_0), and DS_3 (augmented DS_1). We evaluate a range of ML models-Logistic Regression (LR), Decision Tree (DT), Random Forest (RF), Gradient Boosting (GB), XGBoost (XBG), Naive Bayes (NB), Support Vector Machine (SVM), and AdaBoost (AdB)-on these datasets, analyzing prediction accuracy, model size, and prediction latency. The results show that while larger datasets lead to increased model memory footprints and prediction latencies, the Canny edge detection preprocessing supplemented by Hessian filtering (used in DS_1 and DS_3) degrades the performance of most models. In our experiment, we observe that Random Forest (RF) maintains a stable memory footprint of 61 KB across all datasets, while models like KNN and SVM show significant increases in memory usage, from 5.7-7 KB on DS_0 to 102-220 KB on DS_2, and similar increases in prediction time. Logistic Regression, Decision Tree, and Naive Bayes show stable memory footprints and fast prediction times across all datasets. XGBoost's prediction time increases from 180-200 ms on DS_0 to 700-3000 ms on DS_2 (truncated)
Journal refNature Scientific Reports volume 15, Article number: 16413 (2025)
DOI:10.1038/s41598-025-98356-7