arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Khondo:孟加拉语表格文档分组拆分的多模态基准测试

Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms

Abu Tyeb Azad, Fahim Ahmed, Ishita Sur Apan, Ezharuddin Jubaer, Sumaiya Karim Katha, Armun Alam, Amin Ahsan Ali, Aman Chadha, Md Mofijul Islam, AKM Mahbubur Rahman

arXiv 2607.21780首次发表:更新:

发表机构

Wichita State University; CCDS, Independent University, Bangladesh; Amazon GenAI(威奇托州立大学; 孟加拉国独立大学计算与通信科学系; 亚马逊生成式人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对孟加拉语表格文档分组拆分难的问题,引入多模态基准测试Khondo,涵盖多种拼接方案和行政领域。通过对语言模型零样本评估及对照分析,发现页面排列是主要挑战,语言有影响,确立页面顺序重建问题并提供基准测试。

AI 中文摘要

文档包是由多个文档拼接成的单个文件,在政府和行政工作流程中很常见,但将其拆分为单个文档却很困难,尤其是对于低资源语言。我们引入了Khondo(孟加拉语中意为拆分/分割),这是首个针对孟加拉国政府表格文档分组拆分的基准测试。与之前基于英语和OCR文本的数据集不同,Khondo是双语(孟加拉语-英语)且基于视觉原生的,模型直接在页面图像上运行。它涵盖了从顺序到完全混洗的五种拼接方案,跨越14个行政领域,带有真实边界、领域类型和页面顺序。对语言模型的零样本评估表明,它们能较好地将页面聚类到源文档中,但在页面混洗后恢复原始顺序时存在困难。为了找出导致这种困难的原因,我们进行了两项对照分析,分别改变提示指令和文档包语言。两者主要影响排序而非聚类:(a)明确的页面顺序指令是必要但不充分的,(b)英语文档包的排序比孟加拉语更可靠,这使得页面排列成为主要挑战,语言是次要但一致的因素。Khondo将页面顺序重建确立为基于视觉的低资源文档理解中的一个关键开放问题,并为衡量解决该问题的进展提供了一个可控的基准测试。我们的数据集和代码可在这个https网址获取。

英文摘要

Document packets, multiple documents concatenated into a single file, are common in government and administrative workflows, yet splitting them into their constituent documents is difficult, especially for low-resource languages. We introduce Khondo (Bangla for split/segment), the first benchmark for document packet splitting on Bangladeshi government forms. Unlike prior English and OCR-text-based datasets, Khondo is bilingual (Bangla--English) and vision-native; where models operate directly on page images. It spans five concatenation schemes, from sequential to fully shuffled, across 14 administrative domains, with ground-truth boundaries, domain types, and page order. Zero-shot evaluation of MLLMs shows they cluster pages into their source documents fairly well but struggle in restoring the original page order once shuffled. To isolate what drives this difficulty, we run two controlled analyses, varying the prompt instruction and then the packet language. Both primarily affect ordering rather than clustering: (a) explicit page-order instructions are necessary but insufficient, and (b) English packets are ordered more reliably than Bangla, making page arrangement the dominant challenge and language a secondary but consistent factor. Khondo establishes page-order reconstruction as a key open problem in vision-based, low-resource document understanding, and provides a controlled benchmark for measuring progress toward solving it. Our dataset and code is available at https://huggingface.co/datasets/Mausul/khondo

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑