arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BioSentinel在EXIST 2026中的应用:使用XLM-RoBERTa进行软标签优化以实现表情包中的性别歧视意图分类

BioSentinel at EXIST 2026: Soft-Label Optimization with XLM-RoBERTa for Sexism Intent Classification in Memes

Chandru Munisamy, Karthikeya Raguveer, Alapan Kuila

arXiv 2607.24137首次发表:更新:

发表机构

Indian Institute of Information Technology, Design and Manufacturing, Kurnool(印度库努尔信息技术、设计与制造学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BioSentinel团队参与EXIST 2026任务2.2,用基于xlm-roberta-base的文本中心方法,结合复合损失函数训练,在表情包性别歧视意图分类任务中,给出硬标签和软标签预测,通过消融等分析得出相关结论,在测试集取得一定排名。

AI 中文摘要

本文描述了BioSentinel团队参与EXIST 2026任务2.2:表情包中的来源意图,这是CLEF 2026评估活动的一部分。该任务要求在学习不一致(Le-Wi-Di)范式下,将表情包背后的交流意图分类为直接、评判性或无(非性别歧视),该范式要求硬标签和软标签(概率分布)预测。我们提出了一种以文本为中心的方法,基于xlm-roberta-base(2.7亿参数)训练,使用复合损失函数,结合软注释器分布上的KL散度和硬标签上的加权交叉熵。在官方测试集上,该系统实现了ICM-Soft-Norm为0.3229和ICM-Norm为0.3778,硬F1分数为0.4236,在软-软评估中排名第40(共118份提交),在硬-硬评估中排名第49(共187份提交)。我们对数据集特征、探索性更大架构运行以及注释器分歧在主观NLP任务模型设计中的作用进行了分析。消融结果表明,KL损失提高了软标签指标,而CE损失提高了硬标签准确性。我们还报告了单独的验证集温度分析。

英文摘要

This paper describes the BioSentinel team's participation in EXIST 2026 Task 2.2: Source Intention in Memes, part of the CLEF 2026 evaluation campaign. The task requires classifying the communicative intent behind memes as direct, judgemental, or no (non-sexist), under a Learning with Disagreement (Le-Wi-Di) paradigm that mandates both hard-label and soft-label (probability distribution) predictions. We present a text-centric approach built on xlm-roberta-base (270M parameters) trained with a composite loss function combining KL divergence on soft annotator distributions and weighted cross-entropy on hard labels. On the official test set, the system achieved an ICM-Soft-Norm of 0.3229 and ICM-Norm of 0.3778, with a hard F1-score of 0.4236, ranking 40th (out of 118 submissions) in the soft-soft evaluation and 49th (out of 187 submissions) in the hard-hard evaluation. We provide an analysis of the dataset characteristics, exploratory larger-architecture runs, and the role of annotator disagreement in shaping model design for subjective NLP tasks. Ablation results show that KL loss improves soft-label metrics, while CE loss improves hard-label accuracy. We also report a separate validation-set temperature analysis.

Comments9 pages, 1 figure, 6 tables. Accepted for publication in the CLEF 2026 Working Notes (EXIST 2026), Jena, Germany

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑