发表机构
Daffodil International University; Kennesaw State University(达福德国际大学; 肯尼索州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文构建了包含10,000个孟加拉语句子的四类功能标注语料库,通过TF-IDF特征与双层集成模型(DLE)实现0.95的准确率和宏F1,并利用LIME提供可解释性分析。
AI 中文摘要
自动句子功能识别对于许多下游自然语言处理(NLP)应用(如对话系统、文本转语音合成和机器翻译)至关重要。然而,孟加拉语句子功能分类的基准资源仍然有限。为弥补这一不足,本文构建了一个包含10,000个孟加拉语句子的语料库,这些句子被人工标注为四种功能类别,即陈述句、疑问句、祈使句和感叹句。该语料库在四个类别之间几乎平衡,Fleiss' Kappa系数为0.82,反映了较高的标注可靠性。此外,我们评估了多种特征表示,包括词袋(BoW)、TF-IDF和Word2Vec,并结合多种经典机器学习分类器。同时,我们采用了两种异构集成模型,即单层集成(SLE)和双层集成(DLE),以提高分类性能。实验结果表明,TF-IDF始终优于Word2Vec,这可能是由于其能够强调与句子功能相关的判别性词汇线索,尤其是在用于训练Word2Vec的语料库相对较小的情况下。使用TF-IDF特征的DLE模型取得了最佳性能,准确率和宏F1分数均为0.95,证明了稀疏词汇表示和异构集成学习在此任务中的有效性。进一步的交叉验证确认了该方法的稳健性,而基于LIME的可解释性分析为模型预测提供了见解。所开发的语料库和模型基准测试为孟加拉语句子功能分类建立了强有力的基线。
英文摘要
Automatic sentence function identification is important for many downstream natural language processing (NLP) applications such as dialogue systems, text-to-speech synthesis, and machine translation. However, benchmark resources for Bangla sentence function classification remain limited. To mitigate this gap, this paper introduces a corpus of 10,000 Bangla sentences, manually annotated into four functional categories, namely declarative, interrogative, imperative, and exclamatory. The corpus is nearly balanced across the four classes, with high annotation reliability reflected by a Fleiss\' Kappa of 0.82. Furthermore, we evaluate multiple feature representations, including Bag-of-Words (BoW), TF-IDF, and Word2Vec, with several classical machine learning classifiers. In addition, two heterogeneous ensemble models, namely Single-Level Ensemble (SLE) and Double-Level Ensemble (DLE), are utilized to improve classification performance. Experimental results show that TF-IDF consistently outperforms Word2Vec, likely due to its ability to emphasize discriminative lexical cues associated with sentence functions, particularly given the relatively small corpus used to train Word2Vec. The DLE model with TF-IDF features achieves the best performance with accuracy and macro-F1 of 0.95, demonstrating the effectiveness of sparse lexical representations and heterogeneous ensemble learning for this task. Further cross-validation confirms the robustness of the approach, while LIME-based interpretability provides insights into model predictions. The developed corpus and model benchmarking establish strong baselines for Bangla sentence function classification.
CommentsAccepted at 2026 IEEE 5th International Conference on Robotics, Automation, Artificial-Intelligence and Internet-of-Things (RAAICON)