AI 中文总结
针对低资源语言情感分析难题,提出结合领域自适应持续预训练与LoRA参数高效微调的SentiBanglaBERT框架,集成SHAP可解释性,实现高效且可解释的孟加拉语情感分类,性能可媲美强基准。
AI 中文摘要
理解低资源语言的情感分析仍是自然语言处理(NLP)领域的关键挑战,尤其是在领域特定数据稀缺时。本研究提出SentiBanglaBERT,这是一种结合领域自适应持续预训练与参数高效微调的两阶段孟加拉语情感分类框架。该方法可实现对新闻风格数据的上下文自适应,同时通过Low-Rank Adaptation(LoRA)保持计算效率。除性能外,SentiBanglaBERT还集成了基于SHAP的可解释性,为孟加拉语形态线索(如否定后缀和体标记)如何影响情感预测提供语言学见解。实验表明,其稳定性能可与强基准模型媲美,同时具备更高的透明度和解释深度。该框架凸显了领域自适应持续学习作为形态丰富、代表性不足语言的可解释、资源高效NLP基础的潜力。
英文摘要
Understanding sentiment in low-resource languages remains a key challenge for Natural Language Processing (NLP), particularly when domain-specific data is scarce. In this work, we present SentiBanglaBERT, a two-stage Bengali sentiment classification framework combining domain-adaptive continual pretraining and parameter-efficient fine-tuning. The approach enables contextual adaptation to news-style data while remaining computationally efficient through Low-Rank Adaptation (LoRA). Beyond performance, SentiBanglaBERT integrates SHAP-based interpretability, offering linguistic insights into how Bengali morphological cues, such as negation suffixes and aspectual markers, influence sentiment predictions. Experiments demonstrate stable performance comparable to strong baselines while providing greater transparency and interpretive depth. This framework highlights the potential of domain-adaptive continual learning as a foundation for interpretable, resource-efficient NLP in morphologically rich, underrepresented languages.
DOI:10.1109/iSAI-NLP66160.2025.11320544