arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15608cs.CLcs.AIcs.CVcs.LG

透过观察者之眼:面向多模态性别歧视检测的生物特征与人口统计条件化

Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection

Ana-Maria Luisa Mocanu, Sebastian Mocanu, Ciprian-Octavian Truică, Elena-Simona Apostol

首次发表
浏览论文内容

中文总结 AI 辅助

针对互联网性别歧视检测的主观性,提出融合标注者心理与人口统计特征的多模态框架,通过交叉注意力和条件化提升检测性能,在EXIST 2026任务中排名第29。

中文摘要 AI 辅助

检测互联网上的性别歧视本质上是一项主观任务;我们的团队VANGUARD在EXIST 2026任务2中通过提出一个以人为中心的多模态框架来应对这一挑战,该框架分析并将人类标注者的心理和人口统计特征纳入检测流程。我们通过交叉注意力架构与特征级线性调制条件化融合五种输入模态。模因文本使用Gemma 4提取和视觉描述,然后通过NLLB-200在英语和西班牙语之间进行自动翻译增强。文本和图像表示由LoRA适配的XLM-RoBERTa和CLIP编码器生成,并与预训练自编码器编码的传感器特征融合。为建模标注者主观性,我们将子任务2.1构建为标签分布学习问题,优化整个标注者标签分布上的Kullback-Leibler散度损失。在推理时,通过深度多模态网络与在文体特征和生理特征上训练的互补SVM之间的软投票产生预测。我们最好的提交在子任务2.2(源意图)的软评估下在114个队伍中排名第29,并且归一化ICM分数在子任务2.1和2.2上均高于基线,表明以标注者为中心的条件化贡献了可用信号。我们发布完整的流程和分析以支持可复现的以人为中心的建模。

英文摘要

Detecting sexism on the internet is a fundamentally subjective task; our team, VANGUARD, addresses this challenge in the EXIST 2026 Task 2 by proposing a human-centered multimodal framework that analyses and incorporates the psychological and demographic characteristics of human annotators into the detection pipeline. We fuse five input modalities through a cross-attention architecture with Feature-wise Linear Modulation conditioning. Meme text is extracted and visually described with Gemma 4, then augmented by automatic translation between English and Spanish with NLLB-200. Text and image representations are produced by LoRAadapted XLM-RoBERTa and CLIP encoders and fused with sensor features encoded by a pretrained autoencoder. To model annotator subjectivity, we frame Subtask 2.1 as a label distribution learning problem, optimizing a Kullback-Leibler divergence loss over the full annotator label distribution. At inference time, predictions are produced by soft-voting between the deep multimodal network and a complementary SVM trained on stylometric and physiological features. Our best submission ranks 29th out of 114 on Subtask 2.2 (source intention) under soft evaluation, and the normalized ICM scores remain above the baseline on Subtasks 2.1 and 2.2, indicating that annotator-centered conditioning contributes a usable signal. We release our full pipeline and analysis to support reproducible human-centered modeling.

发表机构

  • National University of Science and Technology POLITEHNICA Bucharest(布加勒斯特理工大学)
  • Academy of Romanian Scientists(罗马尼亚科学家学院)

机构由 AI 辅助整理,请以论文原文为准。

↑