MUSE:在情境化教育中的多模态理解上对大型视觉-语言模型进行基准测试
MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
浏览论文内容
中文总结 AI 辅助
提出MUSE基准,评估大型视觉-语言模型在情境化教育中对艺术图像的多模态理解,涵盖十二项任务,揭示模型在情感解读和组合推理上的显著差距,并识别常见失败模式。
中文摘要 AI 辅助
大型视觉-语言模型在多模态理解方面取得了显著进展,但其在教育场景中的能力仍未得到充分评估。在人工智能辅助语言学习中,模型必须解读艺术图像,理解其语义、情感和文化内容,并推理视觉情境以支持有意义的交互。然而,现有基准主要聚焦于真实世界图像或特定领域的教育推理,对艺术教育内容的覆盖有限。为弥补这一空白,我们提出了MUSE,一个用于评估大型视觉-语言模型在情境化教育应用中艺术图像理解能力的基准。MUSE将图像标注与问题生成解耦,使得能够以可控难度生成多样化任务,同时减少标注工作量。它包含十二项任务,涵盖视觉感知、语义与情感解读、文化理解以及组合推理,并配以多样化的艺术图像,这些图像经过精心策划,以新加坡和东南亚多元文化背景为中心,同时兼顾西方艺术传统,覆盖多个主题和难度级别。对开源和专有模型的评估揭示了各能力维度上的显著差异,尤其在情感解读和组合推理方面。我们的分析进一步识别了常见的失败模式,以及为教育开发可信赖多模态模型的关键挑战。我们希望MUSE能作为推进情境化教育应用中多模态理解的标准基准。
英文摘要
Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated. In AI-assisted language learning, models must interpret artistic imagery, understand its semantic, affective, and cultural content, and reason about visual context to support meaningful interaction. However, existing benchmarks primarily focus on real-world images or domain-specific educational reasoning, providing limited coverage of artistic educational content. To address this gap, we introduce MUSE, a benchmark for evaluating large vision-language models on artistic image understanding in situated educational applications. MUSE decouples image annotation from question generation, enabling diverse tasks with controllable difficulty while reducing annotation effort. It comprises twelve tasks spanning visual perception, semantic and affective interpretation, culture understanding, and compositional reasoning, together with diverse artistic images deliberately curated to center Singaporean and Southeast Asian multicultural contexts alongside Western art traditions, covering multiple themes and difficulty levels. Evaluation of open-source and proprietary models reveals substantial disparities across capability dimensions, particularly in affective interpretation and compositional reasoning. Our analysis further identifies common failure modes and key challenges for developing trustworthy multi-modal models for education. We hope MUSE will serve as a standardized benchmark for advancing multi-modal understanding in situated educational applications.
发表机构
- AI Singapore, National University of Singapore(新加坡国立大学新加坡人工智能)
- School of Computing, National University of Singapore(新加坡国立大学计算学院)
- Institute of Advanced Intelligence and Computing, A*STAR(新加坡科技研究局先进智能与计算研究所)
机构由 AI 辅助整理,请以论文原文为准。