基于文本特征分析的研究论文质量识别
Research Paper Quality Recognition Through Textual Feature Analysis
浏览论文内容
中文总结 AI 辅助
本文提出仅用论文标题和摘要文本特征区分优质与非优质论文的基准,结合多种嵌入技术与分类器,实验显示FastText+SVM准确率达91.12%,为学术诚信工具开发提供支持
中文摘要 AI 辅助
科研工作的质量与可信度决定了知识和创新的形成,然而区分有影响力的高质量研究与存在缺陷的研究仍是一项挑战。本文提出一个基准,仅利用论文标题和摘要的文本特征将研究论文分为两类:优质(高被引)和非优质(被撤回)。我们评估了多种嵌入技术,包括SBERT、Word2Vec、FastText、USE和TF-IDF,结合支持向量机(SVM)、随机森林和神经网络等分类器。本文的贡献包括:(1)超参数透明性;(2)基于t-SNE的特征空间可视化;(3)利用SHAP进行模型可解释性分析;(4)对错误案例的详细研究。实验结果显示,结合SBERT嵌入的神经网络达到87.22%的准确率,而FastText结合SVM达到91.12%的准确率。这些发现凸显了文本信息在评估研究质量中的价值,同时也指出了部署时的伦理考量。本研究有助于开发促进学术诚信的工具,推动可信学术的发展。
英文摘要
Knowledge and innovations are shaped by using the quality and credibility of the scientific research. Yet, distinguishing between impactful, high-quality work and flawed studies remains a challenge. This paper introduces a benchmark for classifying research papers into two categories: good (highly cited) and non-good (retracted), using only textual features from titles and abstracts. We evaluate multiple embedding techniques, including SBERT, Word2Vec, FastText, USE, and TF-IDF, combined with classifiers such as Support Vector Machines (SVM), Random Forests, and Neural Networks. Our contributions include: (1) hyperparameter transparency, (2) feature space visualizations using t-SNE, (3) model interpretability analysis with SHAP, and (4) detailed examination of error cases. Experimental results show that a neural network with SBERT embeddings achieves 87.22\% accuracy, while FastText combined with SVM reaches 91.12\%. These findings highlight the value of textual information in assessing research quality, with ethical considerations for deployment. This work contributes toward the development of academic integrity tools that promote trustworthy scholarship.