NormViz:一个用于在全球文化中锚定多模态推理的基准与框架
NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures
浏览论文内容
中文总结 AI 辅助
针对AI文化理解缺失,提出NormViz基准与训练集,含3,268对跨16国图像,微调后模型准确率相对提升达125%,确立视觉规范理解这一多模态新前沿。
中文摘要 AI 辅助
人工智能系统在全球范围内使用,但它们难以满足文化多样性人群的需求。以往关于文化理解的研究仅在纯文本设置或视觉物品识别(如食物、衣物)上评估人工智能系统。通过当地社会规范对可观察到的视觉行为进行推理的能力,我们称之为视觉规范理解,仍未得到检验。我们引入了NormViz-Bench,一个高质量、经过人工验证的基准,包含覆盖16个国家的3,268对对比图像(共6,536张图像)。每对图像仅在改变图像解读方式的文化相关行为(如物体、属性、空间关系和动作)上有所不同。每张图像被标注为符合、违反或与当地社会规范无关,并且成对评估要求两张图像均被正确分类,从而防止依赖表面的视觉捷径。即使是最强的视觉语言模型,Gemini 3.0 Flash和Qwen2.5 VL 7B,也仅在26.6%和21.6%的图像对上成功,最难以识别违反规范和文化上良性的视觉行为。为弥补这一差距,我们引入了NormViz-Train,一个包含64,000张图像及解释的训练数据集。尽管绝对性能仍然较低(低于30%),在NormViz-Train上进行微调分别将Qwen3-VL 4B和8B的成对准确率相对提高了高达125%和36%,展示了教导模型将视觉感知与文化意义联系起来的前进道路。NormViz-Bench和NormViz-Train共同将视觉规范理解确立为多模态人工智能的一个具有挑战性和重要性的前沿领域。
英文摘要
AI systems are used worldwide, but they struggle to serve the needs of culturally diverse populations. Prior work on cultural understanding evaluates AI systems on text-only settings or on visual artifact recognition (e.g. foods, clothing). The ability to reason about visually observable behaviors through local social norms, which we call visual norm understanding, remains unexamined. We introduce NormViz-Bench, a high quality, human-validated benchmark of 3,268 contrastive image pairs (6,536 images) spanning 16 countries. Each pair varies only in the culturally relevant behavior (e.g., objects, attributes, spatial relations, and actions) that alters how each image is interpreted. Each image is labeled as conforming to, violating, or irrelevant to local social norms, and pair-level evaluation requires both images to be correctly classified, thereby preventing reliance on superficial visual shortcuts. Even the strongest VLMs, Gemini 3.0 Flash and Qwen2.5 VL 7B, succeed on only 26.6% and 21.6% of pairs, struggling most with identifying violating and culturally benign visual behaviors. Towards bridging this, we introduce NormViz-Train, a training dataset of 64k images paired with explanations. Though absolute performance remains low (<30%), finetuning on NormViz-Train improves pair accuracy relatively by up to 125% and 36% Qwen3-VL 4B and 8B respectively, showing a path forward to teach models to connect visual perception to cultural significance. Together, NormViz-Bench and NormViz-Train establish visual norm understanding as a challenging and consequential frontier for multimodal AI.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
- University of California Los Angeles(加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。