发表机构
Chittagong University of Engineering & Technology (CUET)(吉大港工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对多模态厌女症检测问题,提出GeoMVC方法,通过几何交互层建模跨模态对齐,用多视图共识策略减轻分布偏移,在相关挑战中取得一定排名成绩,凸显特定文化背景下建模的挑战。
AI 中文摘要
互联网模因的激增给自动内容审核带来新复杂性,尤其是检测厌女症。模因常依赖视觉和文本模态间语义冲突,恶意意图隐含且有文化背景。本文介绍为ICMI 2026的CC-MMD大挑战开发的GeoMVC。针对静态特征拼接局限,提出几何交互层,通过哈达玛积和冻结视觉与文本嵌入间余弦相似度建模跨模态对齐。还通过多视图共识策略减轻噪声OCR和代码混合音译导致的分布偏移,聚合原始、长度过滤和英文翻译文本视图的预测。该系统在任务A的马拉雅拉姆语分区中排名第2(宏F1:0.892),中文分区中排名第3(宏F1:0.895),泰米尔语分区中排名第5(宏F1:0.521)。对开发分区的详细错误分析突出了在跨达罗毗荼语和中文文化背景下对局部音译和代码混合讽刺进行建模的开放挑战。
英文摘要
The proliferation of internet memes has introduced new complexities to automated content moderation, particularly in detecting misogyny. Memes often rely on a semantic clash between visual and textual modalities, where hateful intent is implicit and culturally grounded. This paper presents GeoMVC (Geometric Interaction and Multi-View Consensus), developed for the CC-MMD Grand Challenge at ICMI 2026. To address the limitations of static feature concatenation, a Geometric Interaction Layer is proposed that models cross-modal alignment via Hadamard products and cosine similarity between frozen visual and textual embeddings. We further mitigate distribution shifts caused by noisy OCR and code-mixed transliteration through a Multi-View Consensus strategy, aggregating predictions across raw, length-filtered, and English-translated text views. The system achieved Rank 2 in the Malayalam partition (Macro F1: 0.892) and Rank 3 in the Chinese partition (Macro F1: 0.895) on Task A, while securing Rank 5 in the Tamil partition (Macro F1: 0.521). A detailed error analysis on the development partition highlights open challenges in modeling localized transliteration and code-mixed sarcasm across Dravidian and Chinese cultural contexts.