顺序很重要:面向视觉-语言推理的中文多面板梗图基准
Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning
- Shanghai Maritime University(上海海事大学)
- Dalian Maritime University(大连海事大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究推出中文多面板梗图基准CMPM,评估大型视觉-语言模型对梗图的顺序感知推理能力,发现模型在打乱顺序时准确率大幅下降,Gemini 3.1 Pro和GPT-5.5表现优于开源模型。
AI中文摘要:
许多多模态任务不仅依赖于对视觉元素的孤立识别,还取决于视觉元素的排序与组合方式。互联网梗图是该问题的典型案例:其笑点往往依赖于受限的阅读顺序以及跨面板的视觉-文本线索。尽管大型视觉-语言模型(LVLM)在单图像理解上表现出色,但它们是否能对结构化的梗图布局执行感知顺序的推理,尤其是在中文社交媒体场景下,仍不明确。我们推出了CMPM(Chinese Multi-Panel Meme,中文多面板梗图基准),该基准包含1214个标注样本,涵盖五种结构类型、排序依赖、面板顺序约束以及可选的评论上下文。我们设计了两层评估:任务1探究结构分类与感知顺序的面板排序(含上下文消融设置),任务2对中文梗图解释生成进行评估,采用人类对五个1-3级李克特维度(视觉、面板、幽默、上下文和忠实度)的评分。我们在统一协议下对五个代表性LVLM进行了基准测试。结果表明,规范显示准确率本身并非顺序理解的证据:主要的打乱顺序条件下准确率大幅下降,揭示了感知顺序的多模态推理仍存在持续差距。任务2的偏好显示Gemini 3.1 Pro和GPT-5.5优于开源模型,而评论上下文仅产生小且混杂的Core4增益。代码和数据将在论文接收后发布。
英文摘要:
Many multimodal tasks depend on how visual elements are ordered and composed, not only on recognizing them in isolation. Internet memes are a compact case of this problem: their punchline often depends on a constrained reading order and cross-panel visual--textual cues. While large vision-language models (LVLMs) show strong performance on single-image understanding, it remains unclear whether they can perform sequence-aware reasoning over structured meme layouts, especially in Chinese social media. We introduce CMPM, a Chinese Multi-Panel Meme benchmark with 1,214 annotated samples covering five structural types, ordering dependency, panel-order constraints, and optional comment context. We formulate a two-layer evaluation: Task1 probes structure typing and order-sensitive panel sequencing (with a context ablation setting), and Task2 evaluates Chinese meme explanation generation with human ratings on five 1-3 Likert dimensions (visual, panel, humor, context, and faithfulness). We benchmark five representative LVLMs under a unified protocol. Results indicate that canonical-display accuracy is not by itself evidence of order understanding: the primary shuffled condition produces a sharp accuracy drop, revealing a persistent gap in order-sensitive multimodal reasoning. Task2 preferences place Gemini 3.1 Pro and GPT-5.5 above the open models, while comment context yields only a small and mixed Core4 gain. Code and data will be released upon acceptance.