发表机构
Old Dominion University(奥多明尼昂大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估了三种定制工具和三种LLM与人类在100条推文上的情绪分析一致性,发现人类间仅一般一致,Twitter-roBERTa-base与人类最对齐,强调领域微调和以人为中心评估的重要性。
AI 中文摘要
社交媒体是实时公众情绪的丰富来源,但广泛使用的情绪分析工具往往在未了解其局限性的情况下被应用。在本研究中,我们评估了三种定制情绪分析工具(TextBlob、VADER和Twitter-roBERTa-base)和三种大型语言模型(LLM:Qwen3-32B、GPT-OSS-120B、Llama-4-Maverick-17B)与六名人类评分者在100条推文上的评分者间信度。我们使用两种统计度量来衡量一致性:用于两两比较的Cohen's kappa和用于多个评分者的Fleiss' kappa。即使在人类评分者之间,我们的结果也仅显示出一般的一致性,凸显了情绪分析的主观性。在二元情绪分类(负面与非负面,正面与非正面)下观察到的一致性高于三分类,这在人类和自动化工具中均如此。Twitter-roBERTa-base模型与人类评分的一致性最强,优于定制情绪工具和LLM,尤其在区分负面与非负面情绪方面。LLM之间表现出显著的一致性,并与人类有中度至显著的一致性,在正面与非正面分类中表现更好。我们的发现强调,针对特定领域的微调对于可靠的社交媒体情绪分析仍然至关重要,而以人为中心的评估对于建立金标准标签仍然必不可少。
英文摘要
Social media is a rich source of real-time public sentiment, but widely used sentiment analysis tools are often applied without understanding their limitations. In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets. We measured agreement using two statistical measures: Cohen's kappa for pairwise comparisons and Fleiss' kappa for multiple raters. Even among the human raters, our results showed only fair agreement, highlighting the subjectivity of sentiment analysis. Higher agreement was observed under the binary sentiment classification (negative vs. non-negative and positive vs. non-positive) than under the three-class classification across both humans and automated tools. The Twitter-roBERTa-base model showed the strongest alignment with human ratings, outperforming both bespoke sentiment tools and LLMs, particularly in distinguishing negative versus non-negative sentiment. LLMs showed substantial agreement among themselves and moderate to substantial alignment with humans, performing better in positive vs. non-positive classifications. Our findings underscore that domain-specific fine-tuning remains crucial for reliable social media sentiment analysis, and human-centered evaluation remains essential for establishing gold-standard labels.
Comments11 pages, 1 figure, 2 tables, accepted for publication at TPDL 2026