发表机构
SRM Institute of Science and Technology; Amity University Dubai; University of Sharjah(SRM科学技术学院; 阿米提大学迪拜分校; 沙迦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用UCI心脏病数据集,对比传统机器学习与LLM生成规则系统,发现传统模型性能更优(随机森林准确率90.2%),但LLM规则提供可解释的IF-THEN逻辑,凸显了预测性能与可解释性之间的权衡。
AI 中文摘要
本研究使用UCI心脏病数据集,比较了传统机器学习模型与大语言模型(LLM)生成的基于规则的系统在心脏病预测中的表现。研究评估了包括逻辑回归、K近邻(KNN)、支持向量机(SVM)、朴素贝叶斯、决策树和随机森林在内的多种分类器,以及使用GPT-4o和Claude Sonnet 4.6生成的基于规则的系统。模型性能通过准确率、精确率、召回率和F1分数进行评估。实验结果表明,传统机器学习模型在预测性能上始终优于LLM生成的基于规则的系统。随机森林取得了最佳整体性能,准确率为90.2%,精确率为0.829,召回率完美达到1.0,F1分数为0.906。朴素贝叶斯紧随其后,准确率为88.5%,F1分数为0.881。相比之下,LLM生成的规则模型性能较低,其中Claude Sonnet 4.6达到80.3%的准确率(F1分数:0.833),GPT-4o获得70.5%的准确率(F1分数:0.690)。尽管存在性能差距,LLM生成的规则提供了可解释的IF-THEN诊断逻辑,增强了临床决策中的可解释性和透明度。这些发现突显了医疗人工智能系统中预测性能与可解释性之间的权衡。所有实验的完整实现,包括机器学习模型和LLM派生的规则分类器,均可在GitHub仓库中公开获取,链接见本https URL。
英文摘要
This study compares traditional machine learning models and Large Language Model (LLM)-generated rule-based systems for heart disease prediction using the UCI Heart Disease dataset. Several classifiers, including Logistic Regression, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Naive Bayes, Decision Tree, and Random Forest, were evaluated alongside rule-based systems generated using GPT-4o and Claude Sonnet 4.6. Model performance was assessed using accuracy, precision, recall, and F1-score metrics. Experimental results show that traditional machine learning models consistently outperform LLM-generated rule-based systems in predictive performance. Random Forest achieved the best overall performance with 90.2% accuracy, a precision of 0.829, perfect recall of 1.0, and an F1-score of 0.906. Naive Bayes followed closely with 88.5% accuracy and an F1-score of 0.881. In contrast, the LLM-generated rule models achieved lower performance, with Claude Sonnet 4.6 reaching 80.3% accuracy (F1-score: 0.833) and GPT-4o obtaining 70.5% accuracy (F1-score: 0.690). Despite the performance gap, the LLM-generated rules provide interpretable IF-THEN diagnostic logic that enhances explainability and transparency in clinical decision-making. These findings highlight the trade-off between predictive performance and interpretability in medical artificial intelligence systems. The complete implementation of all experiments, including machine learning models and LLM-derived rule classifiers, is publicly available in the GitHub repository at https://github.com/FeisalAlaswad/LLM-Rule-ML-Heart-Disease-Prediction .