Beyond the Literal: Decomposing Pragmatic Intent in Multimodal Meme Understanding
超越字面:多模态模因理解中的语用意图分解
机构 * The Chinese University of Hong Kong(香港中文大学) ; Huawei(华为) ; Central China Normal University(中央师范大学) ; Shenzhen University(深圳大学) ; King’s College London(伦敦国王学院) ; University of International Relations(国际关系大学)
AI总结 针对大型视觉语言模型(LVLMs)在理解模因时倾向于描述字面内容而非语用意图的问题,提出Intent Projection框架,通过表示、输出和目标三层面的字面-语用分解,在六个基准上超越开源模型并缩小与专有模型的差距。