发表机构
Delhi Technological University; Multimodal Data Analytics Research Laboratory(德里理工大学; 多模态数据分析研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态讽刺和网络欺凌检测难题,提出HCIG框架,利用图注意力网络在不同层面建模跨模态不协调并整合,引入GCCN辅助推理。实验表明该方法在相关数据集上表现出色,分层多粒度建模比传统策略更有效。
AI 中文摘要
多模态讽刺和网络欺凌检测具有挑战性,因为其含义常源于文本和视觉信息的不协调。现有方法主要依赖特征融合或跨模态注意力,难以捕捉不同层次表示的语义不一致。本文提出HCIG框架,用图注意力网络在词元、短语和全局层面建模跨模态不协调,并通过分层注意力机制整合。还引入GCCN进行基于图的推理。在MMSD讽刺基准和MultiBully网络欺凌数据集上评估,结果表明HCIG在MMSD上性能最佳,GCCN在MultiBully上宏F1最高,分层多粒度不协调建模比传统融合策略更有效。
英文摘要
Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from incongruity between textual and visual information rather than from either modality alone. Existing multimodal approaches primarily rely on feature fusion or cross-modal attention, which may not effectively capture hierarchical semantic inconsistencies across different levels of representation. To address this limitation, this paper proposes HCIG (Hierarchical Cross-modal Incongruity Graph Network), a novel framework that models cross-modal incongruity at token, phrase, and global levels using graph attention networks and adaptively integrates these representations through a learned hierarchical attention mechanism. As a complementary architecture, we also introduce GCCN (Graph-based Cross-modal Contradiction Network), which performs graph-based reasoning using contradiction-aware pooling for efficient multimodal interaction learning. The proposed models are evaluated on the MMSD sarcasm benchmark and the MultiBully cyberbullying dataset, together with comprehensive ablation studies and cross-task transfer experiments. Experimental results demonstrate that HCIG achieves the best performance on MMSD with 85.74% accuracy and 85.29% macro-F1, while GCCN attains the highest macro-F1 (68.66%) on MultiBully and HCIG achieves the highest accuracy (69.62%) and bullying-class F1 (74.90%). The findings demonstrate that hierarchical multi-granularity incongruity modeling provides more effective multimodal reasoning than conventional fusion strategies, offering a robust framework for sarcasm and cyberbullying detection in social media.