发表机构
Yeshiva University(叶史瓦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较了传统机器学习、Transformer模型及少样本提示方法在医疗转录文本临床领域分类中的表现,发现数据平衡显著提升性能,其中BERT在SMOTE平衡数据上取得最高F1分数0.996。
AI 中文摘要
临床领域分类在组织和分析大量非结构化医疗文本中发挥着重要作用。然而,医疗转录数据集往往高度不平衡,这会显著降低分类性能,尤其是对代表性不足的临床专科。在本工作中,我们提出了一项关于基于医疗转录文本的临床领域分类的机器学习与基于Transformer的方法的比较研究。我们评估了六种传统机器学习分类器——朴素贝叶斯、支持向量机(SVM)、决策树、随机森林、K近邻(KNN)和XGBoost——以及两种预训练Transformer模型BERT和XLNet,并采用了一种少样本大语言模型提示方法。实验在从MTSamples收集的医疗转录数据上进行,该数据集包含40个临床专科的5,013个样本。为解决严重的类别不平衡问题,我们研究了两种平衡策略:基于NLP的同义词替换文本增强和合成少数类过采样技术(SMOTE)。实验结果表明,数据平衡显著提升了所评估模型的分类性能。特别是,BERT在SMOTE平衡数据集上取得了最高的F1分数0.996,同时训练时间低于XLNet。这些结果凸显了基于Transformer的表示结合适当的数据平衡策略在临床领域分类中的有效性,并为医学转录分析提供了经典机器学习、Transformer模型和少样本提示的系统比较。
英文摘要
Clinical domain classification plays an important role in organizing and analyzing large volumes of unstructured medical text. However, medical transcription datasets are often highly imbalanced, which can substantially degrade classification performance, particularly for underrepresented clinical specialties. In this work, we present a comparative study of machine learning and transformer-based approaches for clinical domain classification from medical transcriptions. We evaluate six traditional machine learning classifiers---Naive Bayes, Support Vector Machine (SVM), Decision Tree, Random Forest, K-Nearest Neighbors (KNN), and XGBoost---along with two pretrained transformer models, BERT and XLNet, and a few-shot large language model prompting approach. Experiments are conducted on medical transcription data collected from MTSamples, comprising 5,013 samples across 40 clinical specialties. To address severe class imbalance, we investigate two balancing strategies: text augmentation using NLP-based synonym replacement and Synthetic Minority Over-sampling Technique (SMOTE). Experimental results demonstrate that data balancing substantially improves classification performance across the evaluated models. In particular, BERT achieves the highest F1-score of 0.996 on the SMOTE-balanced dataset while requiring lower training time than XLNet. The results highlight the effectiveness of transformer-based representations combined with appropriate data balancing strategies for clinical domain classification and provide a systematic comparison of classical machine learning, transformer models, and few-shot prompting for medical transcription analysis.