发表机构
School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对有害模因检测,提出增强视觉语言多CoT(EVL-MCoT)方法,通过促进多CoT及设计解码框架,解决现有方法局限,在相关数据集上获良好结果,且公开了源代码。
AI 中文摘要
模因在互联网上广泛使用,理解其隐藏含义通常需要文本和视觉的联合解释。现有方法缺乏背景信息和先验知识。可行的选择是采用思维链(CoT),但简单的CoT方法缺乏多视角思维,依赖浅特征融合。本文提出增强视觉语言多CoT(EVL-MCoT)方法,促进多CoT,设计解码框架,在HatefulMemes和MultiOff数据集上取得良好结果,源代码已公开。
英文摘要
MEMEs are widely used on the internet and often carry strong elements of sarcasm or irony. Understanding their hidden meanings typically requires a joint interpretation of text and vision. Existing methods focus on the dual-stream vision-language model to extract the visual and text simultaneously, which lacks background information and prior knowledge about the comprehensive explanation of MEME. One feasible option is to adopt chain-of-thought (CoT). However, the simple CoT approach lacks multi-perspective thinking, which may compromise the reliability of the resulting answers. Moreover, it often relies on shallow feature fusion, lacking the fusion of local details and fine-grained visual-prompt text alignment. This limitation prevents a deeper understanding of the intricate connections between the visual and the text. Herein, an enhanced vision-language multi-CoT (EVL-MCoT) approach is proposed to address these limitations. By promoting multi-CoT, EVL-MCoT enhances consistency and reduces bias in the decision-making process. Additionally, we design a prototype-guided and context-guided decoding framework, which incorporates visual prototypes to guide the fusion process and enables the model to align textual and visual information more precisely. We achieve promising results on the HatefulMemes and MultiOff datasets. The source code has been publicly released and is available at https://github.com/BGWH123/EVL-MCoT.
DOI:10.1007/978-981-95-3349-7_31