arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态大语言模型(MLLMs)真的理解低资源高棉语文档吗?一项关于高棉语文档视觉问答(VQA)的试点研究

Do MLLMs Really Understand Low-Resource Khmer Documents? A Pilot Study on Khmer Document VQA

Nimol Thuon, Panhapin Theang

arXiv 2608.28635首次发表:更新:

发表机构

University of Science and Technology of China; Université Paris Cité(中国科学技术大学; 巴黎城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对高棉语文档VQA,评估开源MLLMs的性能,发现直接使用Qwen3-VL-8B准确率51.9%,外部OCR辅助的Tesseract、PaddleOCR分别达61.9%、61.6%,但高棉语文档理解仍存挑战。

AI 中文摘要

近期,多模态大语言模型(MLLMs)在文档理解、视觉问答(VQA)和文本提取方面取得了进展,但其在低资源非拉丁语境下的可靠性仍不确定。高棉语格式文档存在特殊挑战,包含复杂字体形式、高棉语-英语混合字段、柬埔寨瑞尔与美元两种货币值,且现有高棉语文档VQA资源有限。本文对开源MLLMs开展试点诊断评估,针对高棉语文档图像,从此前推出的KH-FUNSD集合中构建评估子集,涵盖发票、收据、报价单及其他业务表单,该子集包含英语和高棉语问题,答案保留原英语、高棉语、混合字体或数字形式。本研究未推出完整公开基准,而是探究现有模型的能力与失败模式,使用直接图像提示评估代表性开源Qwen-VL模型,并对比解析器辅助及外部OCR辅助配置下的Qwen3-VL-8B性能。直接使用Qwen3-VL-8B的整体准确率达51.9%,优于更小的模型,但高棉语和混合字体答案的性能仍有限;外部OCR表现最佳,Tesseract达61.9%,PaddleOCR达61.6%,不过高棉语答案仍比英语和数字字段难得多。结果表明,当前MLLMs可处理视觉清晰的英语和结构化数字内容,但可靠的原生高棉语文档理解仍是未解决的挑战。

英文摘要

Recent multimodal large language models (MLLMs) have advanced document understanding, visual question answering, and text extraction. However, their reliability in low-resource, non-Latin settings remains uncertain. Khmer form documents present particular challenges because they contain complex script forms, mixed Khmer-English fields, and monetary values in both Cambodian Riel and US Dollars. Available resources for Khmer Document VQA are also limited. This paper presents a pilot diagnostic evaluation of open MLLMs on Khmer document images. We construct an evaluation subset from the previously introduced KH-FUNSD collection, covering invoices, receipts, quotations, and other business forms. The subset includes questions in English and Khmer, with answers retained in their original English, Khmer, mixed-script, or numeric forms. Rather than introducing a full public benchmark, this study examines the capabilities and failure modes of existing models. We evaluate representative open Qwen-VL models using direct image-based prompting and compare parser-assisted and external OCR-assisted configurations with Qwen3-VL-8B. Direct Qwen3-VL-8B outperforms smaller models, achieving 51.9% overall accuracy, although performance remains limited for Khmer-script and mixed-script answers. External OCR produces the strongest results, reaching 61.9% with Tesseract and 61.6% with PaddleOCR. Nevertheless, Khmer-script answers remain substantially more difficult than English and numeric fields. The results indicate that current MLLMs can process visually clear English and structured numeric content, but reliable native Khmer document understanding remains an open challenge.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑