发表机构
Centro de Investigaciones en Computación (CIC), Instituto Politécnico Nacional (IPN)(国家理工学院计算机研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对社交媒体政治观点开展多类别情感分析,设计并评估XGBoost与BERT两种机器学习方法,在测试集上分别取得0.2835与0.2806的F1值,为该领域研究提供了基准。
AI 中文摘要
社交媒体的快速发展催生了海量政治言论,为分析公众意见、识别不同政治视角提供了宝贵机会。情感分析(SA)是自然语言处理(NLP)领域的核心任务,可对文本数据中的态度与观点开展计算研究,对理解政治言论愈发重要。本研究针对社交媒体上的政治观点开展多类别情感分析,即自动区分针对政治议题与人物的多种情感类别。为解决该任务,我们设计并评估了两种基于XGBoost和BERT的机器学习方法。我们使用标准分类指标,在带标注的政治社交媒体帖子数据集上训练并评估模型。实验结果显示,XGBoost模型在测试集上的F1值为0.2835,基于BERT的模型F1值为0.2806。这些结果表明,对复杂且具有上下文的政治言论情感进行分类存在挑战,并为未来多类别政治情感分析研究提供了基准。
英文摘要
The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public opinions and identify different political perspectives. Sentiment Analysis (SA) is a core task in Natural Language Processing (NLP) that allows the computational study of attitudes and opinions in textual data, and has become increasingly important for understanding political discourse. In this work, we investigate multiclass sentiment analysis of political view- points on social media, that is to automatically discriminate multiple sentiment classes over political issues and figures. To solve this task we design and evaluate two machine-learning approaches based on XGBoost and BERT. We train and evaluate the models on a labeled dataset of political social media posts using standard classification metrics. The experimental results show that the XGBoost model reaches an F1-score of 0.2835 and the BERT- based model reaches an F1-score of 0.2806 on the test set. These results demonstrate the challenge of classifying complex and contextualized political discourse sentiment and provide a baseline for future research in multiclass political sentiment analysis.