阿拉伯自然语言处理研究的综合分析:趋势、主题演变与研究缺口——一项基于文献计量和主题的研究
A Comprehensive Analysis of Arabic Natural Language Processing Research: Trends, Topic Evolution, and Research Gaps -- A Bibliometric and Topic-Based Study
浏览论文内容
中文总结 AI 辅助
本研究对7120篇阿拉伯NLP论文开展文献计量与主题分析,明确研究趋势、缺口及引文影响因素,为该领域发展提出针对性建议。
中文摘要 AI 辅助
过去十年,受阿拉伯世界数字化转型、社交媒体以及大语言模型(LLMs)的驱动,自然语言处理(NLP)领域发展迅速。尽管该领域发展迅猛,但目前仍缺乏对其开展的全面定量元分析。本研究对来自6个文献集合的1960年至2026年间发表的7120篇阿拉伯NLP论文,开展了大规模的文献计量及主题分析。研究采用BERTopic进行主题建模,运用回归分析识别引文预测因子,通过社会网络分析研究合著结构,并绘制地理分布图。研究结果显示,2020年后相关论文发表量出现显著激增,驱动因素为Transformer模型与LLMs。主题建模共识别出19个实质性主题,其中最大主题围绕文本、语音、翻译与识别展开。引文分析发现,论文发表时长与引文数呈正相关(相关系数r=0.245,p值小于0.001);回归分析表明,被OpenAlex或Semantic Scholar收录以及机构隶属关系与更高的引文数相关。沙特阿拉伯、美国和埃及的研究产出位居前列。研究构建的任务-方言缺口矩阵识别出关键的未充分研究领域,包括马格里布方言、伊拉克方言和苏丹方言的摘要任务。最大主题的H指数为87,其次是情感分析主题,其H指数为54。本研究的定量方法补充了现有定性调查,并提出了优先研究资源不足的方言、开发符合阿拉伯NLP文化特性的基准等建议。
英文摘要
Arabic Natural Language Processing (NLP) has grown rapidly over the past decade, driven by digital transformation in the Arab world, social media, and large language models (LLMs). Despite this growth, a comprehensive quantitative meta-analysis remains absent. This study presents a bibliometric and topic-based analysis of 7,120 Arabic NLP papers published between 1960 and 2026, sourced from five platforms (arXiv, ACL Anthology, Semantic Scholar, Crossref, OpenAlex) plus an additional targeted OpenAlex subset. We employ BERTopic for topic modeling, regression analysis, social network analysis, and geographic mapping. Our findings show a significant publication surge after 2020, driven by transformer models and LLMs. Topic modeling identifies 19 themes, the largest centered on text, speech, translation, and recognition (2,942 papers). Citation analysis reveals a positive correlation between paper age and citations (r = 0.245, p < 0.001); regression (R^2 = 0.105) shows that indexing in OpenAlex or Semantic Scholar and institutional affiliation are associated with higher citations. Saudi Arabia, the United States, and Egypt lead in research output. A task-dialect gap matrix identifies understudied areas, including summarization for Maghrebi, Iraqi, and Sudanese dialects. The largest topic has the highest H-index (90), followed by sentiment analysis (57). Our quantitative approach complements existing qualitative surveys and offers recommendations to prioritize under-resourced dialects and develop culturally aligned benchmarks.
发表机构
- Kazan Federal University(喀山联邦大学)
机构由 AI 辅助整理,请以论文原文为准。