arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PUMA:面向文化扎根多模态理解的波兰基准测试

PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding

Sławomir Dadas, Michał Perełkiewicz, Rafał Poświata, Małgorzata Grębowiec, Bartłomiej Jaworski, Izabela Woźniakowska

arXiv 2608.21853首次发表:更新:

发表机构

National Information Processing Institute(国家信息处理研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对英语外文化多模态评估不足的问题,提出波兰多模态基准PUMA,通过900项任务评估各类模型在波兰语境下的多模态能力,发现商用模型在视觉问答表现好但音频/文档理解弱,并开源评估框架推进本地化多模态AI研究。

AI 中文摘要

大型语言模型正日益超越文本处理范畴,新增对图像、音频等其他模态的支持。尽管文本理解与生成已得到广泛研究,但多模态数据处理能力,尤其是在英语以外的文化与语言语境下,尚未得到全面评估。本文提出PUMA(Polish Unified Multimodal Assessment,波兰统一多模态评估),这是一个包含900项手工任务的新型基准测试,旨在探究多模态模型在波兰文化与语言语境下的性能极限。该数据集评估了文本、图像、音频及视觉丰富文档处理中的文化理解与实用技能。我们对前沿商用模型、开源权重模型及专用小型系统的广泛评估凸显出显著的性能差距:顶级商用模型在视觉问答中取得高分,但多数模型在复杂音频或文档理解上存在困难。我们开源了评估框架,以推进本地化多模态AI研究。

英文摘要

Large language models are increasingly moving beyond text processing, adding support for other modalities such as images and audio. While text understanding and generation have been extensively studied, multimodal data processing capabilities, particularly in the context of cultures and languages other than English, have not yet been evaluated comprehensively. In this paper, we propose PUMA (Polish Unified Multimodal Assessment), a novel benchmark of 900 hand-crafted tasks designed to probe the limits of multimodal models in the Polish cultural and linguistic context. The dataset evaluates both cultural understanding and practical skill in processing text, images, audio, and visually rich documents. Our extensive evaluation of frontier commercial models, open-weights models, and specialized smaller systems highlights a significant performance gap. While top commercial models achieve high scores in visual question answering, most models struggle with complex audio or document understanding. We open-source our evaluation framework to advance localized multimodal AI research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑