arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31651cs.CVcs.AIcs.CLcs.LG

PalmLeaf-VQA:面向跨区域历史棕榈叶手稿理解的多脚本视觉问答基准

PalmLeaf-VQA: A Multi-Script Visual Question Answering Benchmark for Historical Palm-Leaf Manuscript Understanding Across Diverse Regions

Nimol Thuon, Jun Du, Panhapin Theang

首次发表
浏览论文内容

中文总结 AI 辅助

PalmLeaf-VQA构建了一个包含923幅图像和7,384个问答对的多脚本棕榈叶手稿视觉问答基准,评估多模态大模型在跨区域历史文档理解上的表现,发现最强模型仅达58%准确率,揭示了现有模型在稀有脚本和退化布局理解上的重大局限。

中文摘要 AI 辅助

历史手稿在很大程度上缺席于现代视觉语言基准,这留下了多模态大语言模型(MLLMs)如何处理文化多样、退化且非拉丁文档图像的问题。我们引入了PalmLeaf-VQA,一个面向南亚和东南亚传统中历史棕榈叶手稿理解的多脚本视觉问答基准。PalmLeaf-VQA包含来自八个收藏组的923张精选手稿图像和7,384对问答对:巴厘文、格兰塔文、Jathakam、Kambaramayanam、卡纳达文、高棉文、巽他文和泰米尔文。与面向识别的资源不同,该基准针对与保存相关的线索进行手稿感知的视觉推理,包括物理状况、线条结构、材料和涂层、装订孔、页边距、符号、图画和局部视觉伪影。我们在开放式回答和受限式回答提示下评估了最近的专有和开放权重MLLMs,并跨收藏、问题类别和任务类型提供了细粒度分析。最强的评估模型在保留测试分割上仅达到58.00%的精确匹配准确率,揭示了当前MLLMs在稀有脚本、退化布局和面向保存的文档理解方面的重大局限性。PalmLeaf-VQA为推进文化扎根和布局感知的多模态文档分析提供了一个标准化基准。

英文摘要

Historical manuscripts remain largely absent from modern vision-language benchmarks, leaving open how well multimodal large language models (MLLMs) handle culturally diverse, degraded, and non-Latin document images. We introduce \textbf{PalmLeaf-VQA}, a multi-script visual question answering benchmark for historical palm-leaf manuscript understanding across South and Southeast Asian traditions. PalmLeaf-VQA contains \textbf{923 curated manuscript images} and \textbf{7,384 question--answer pairs} from eight collection groups: Balinese, Grantha, Jathakam, Kambaramayanam, Kannada, Khmer, Sundanese, and Tamil. Unlike recognition-oriented resources, the benchmark targets manuscript-aware visual reasoning over preservation-relevant cues, including physical condition, line structure, material and coating, binding holes, margins, symbols, drawings, and localized visual artifacts. We evaluate recent proprietary and open-weight MLLMs under open-answer and constrained-answer prompting and provide fine-grained analysis across collections, question categories, and task types. The strongest evaluated model reaches only \textbf{58.00\% exact-match accuracy} on the held-out test split, revealing substantial limitations in current MLLMs for rare-script, degraded-layout, and preservation-oriented document understanding. PalmLeaf-VQA provides a standardized benchmark for advancing culturally grounded and layout-aware multimodal document analysis.

↑