发表机构
Manipal University Jaipur(斋浦尔马尼帕尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对讽刺与AI改写社交文本开展情感分析,发现AI改写文本可提升分类器准确率,提出轻量弃权包装器,推动高风险情感应用转向不确定性感知预测。
AI 中文摘要
情感分类器越来越多地被应用于社交媒体内容,这些内容要么是讽刺性的,要么是AI生成的——这两种分布 regime 下,标准评估几乎无法提供指导。我们针对情感分类器在这些场景下的行为开展了一项三部分的实证研究。首先,我们发现讽刺文本的置信度得分显著低于非讽刺文本(Mann-Whitney检验p值为2×10⁻⁶),证实即使没有显式的不确定性建模,分类器也能感知到反讽内容的自身不确定性。其次,与直觉相悖的是,我们发现情感分类器在AI改写的评论上的准确率高于原始人类撰写的文本(RoBERTa模型:针对Qwen3.5-4B改写的文本准确率提升5.8个百分点,针对Gemma4-E4B改写的文本提升3.7个百分点),这揭示了一种跨域风格对齐效应:AI改写消除了混淆Twitter训练分类器的分布噪声,产生了更干净、更具典型性的情感文本。第三,我们证明了一个轻量弃权(不执行)包装器——标记置信度低于0.6的14%输入——在保留集上将准确率从82.2%提升至88.9%(提升6.7个百分点)。我们进一步比较了语义熵和MC-Dropout风格的分歧作为不确定性信号,发现讽刺文本上的AUROC几乎相同(0.650对比0.646),表明对于短社交媒体输入,两种方法可互换。我们的结果推动高风险情感应用(如心理健康标记和内容审核)从置信度单标签预测向不确定性感知弃权(不执行)转变。
英文摘要
Sentiment classifiers are increasingly applied to social media content that is either sarcastic or AI-generated --- two distributional regimes where standard evaluations offer little guidance. We present a three-part empirical study of sentiment classifier behaviour under these conditions. First, we find that confidence scores on sarcastic text are significantly lower than on non-sarcastic text (Mann--Whitney $p = 2 \times 10^{-6}$), confirming that classifiers sense their own uncertainty on ironic content even without explicit uncertainty modelling. Second, and counterintuitively, we show that sentiment classifiers achieve higher accuracy on AI-paraphrased reviews than on the original human-authored text (RoBERTa: $+5.8$ pp for Qwen3.5-4B paraphrases, $+3.7$ pp for Gemma4-E4B), revealing a cross-domain stylistic alignment effect: AI paraphrases remove distributional noise that confounds Twitter-trained classifiers, producing cleaner, more prototypical sentiment text. Third, we demonstrate that a lightweight abstention wrapper --- flagging the $14\%$ of inputs with confidence below $0.6$ --- improves accuracy from 82.2\% to 88.9\% ($+6.7$ pp) on the retained set. We further compare Semantic Entropy and MC-Dropout-style disagreement as uncertainty signals and find near-identical AUROC ($0.650$ vs.\ $0.646$) on sarcastic text, suggesting that for short social media inputs, both methods are interchangeable. Our results motivate a shift from confident single-label prediction to uncertainty-aware abstention in high-stakes sentiment applications such as mental health flagging and content moderation.