arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

两阶段孟加拉语情感分类:通过持续学习和参数高效微调实现领域自适应

Two-Stage Bengali Sentiment Classification: Domain Adaptation Through Continual Learning and Parameter-Efficient Fine-Tuning

MD Shaikh Rahman, Syed Maudud E Rabbi, Muhammad Mahbubur Rashid

arXiv 2608.01471首次发表:更新:

AI 中文总结

针对低资源语言情感分析难题,提出结合领域自适应持续预训练与LoRA参数高效微调的SentiBanglaBERT框架,集成SHAP可解释性,实现高效且可解释的孟加拉语情感分类,性能可媲美强基准。

AI 中文摘要

理解低资源语言的情感分析仍是自然语言处理(NLP)领域的关键挑战,尤其是在领域特定数据稀缺时。本研究提出SentiBanglaBERT,这是一种结合领域自适应持续预训练与参数高效微调的两阶段孟加拉语情感分类框架。该方法可实现对新闻风格数据的上下文自适应,同时通过Low-Rank Adaptation(LoRA)保持计算效率。除性能外,SentiBanglaBERT还集成了基于SHAP的可解释性,为孟加拉语形态线索(如否定后缀和体标记)如何影响情感预测提供语言学见解。实验表明,其稳定性能可与强基准模型媲美,同时具备更高的透明度和解释深度。该框架凸显了领域自适应持续学习作为形态丰富、代表性不足语言的可解释、资源高效NLP基础的潜力。

英文摘要

Understanding sentiment in low-resource languages remains a key challenge for Natural Language Processing (NLP), particularly when domain-specific data is scarce. In this work, we present SentiBanglaBERT, a two-stage Bengali sentiment classification framework combining domain-adaptive continual pretraining and parameter-efficient fine-tuning. The approach enables contextual adaptation to news-style data while remaining computationally efficient through Low-Rank Adaptation (LoRA). Beyond performance, SentiBanglaBERT integrates SHAP-based interpretability, offering linguistic insights into how Bengali morphological cues, such as negation suffixes and aspectual markers, influence sentiment predictions. Experiments demonstrate stable performance comparable to strong baselines while providing greater transparency and interpretive depth. This framework highlights the potential of domain-adaptive continual learning as a foundation for interpretable, resource-efficient NLP in morphologically rich, underrepresented languages.

DOI:10.1109/iSAI-NLP66160.2025.11320544

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑