arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EMBLEM:通过掩码增强多文字表格检测

EMBLEM: Enhancing Multi-script Table Detection through Masking

Dhruv Kudale, Udhay Brahmi, Ganesh Ramakrishnan

arXiv 2609.08330首次发表:更新:

发表机构

Indian Institute of Technology Bombay; BharatGen(印度理工学院孟买分校; 巴拉特基因公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多语言多文字文档表格检测难题,提出掩码范式EMBLEM和数据集MANDALA,仅用英文数据微调即在多文字测试集上提升F1达20.8%。

AI 中文摘要

表格检测是文档分析中的核心任务,支撑着信息检索、文档重建和视觉问答等下游应用。虽然现有的深度学习模型在英文和中文文档上表现良好,但由于文字多样性和标注数据有限,它们在多语言、多文字文档上表现不佳。为了应对这一挑战,我们引入了MANDALA(多文字标注表格检测文档集),这是一个人工整理的数据集,包含来自18种语言、15种文字、涵盖不同领域的2323个含表格页面。我们还提出了EMBLEM,一种基于掩码的多文字表格检测(MTD)范式。EMBLEM生成掩码图像,隐藏文字和字体特有的细节,使在大量英文文档上预训练的模型能够专注于与文字无关的页面布局。在三种表格检测架构上的实验表明,EMBLEM在MANDALA上始终优于强基线,同时在五个标准英文主导的基准上保持竞争力。仅使用英文掩码图像进行微调,不使用任何多文字训练数据,EMBLEM在MANDALA上取得了20.8%的绝对F1分数提升。我们在此https URL发布了MANDALA以及附带的代码和模型。

英文摘要

Table detection is a core task in document analysis, supporting downstream applications such as information retrieval, document reconstruction, and visual question answering. While existing deep learning models perform well on English and Chinese documents, they struggle with multilingual, multi-script documents due to script diversity and the limited availability of labeled data. To address this challenge, we introduce MANDALA (Multi-script Annotated Documents for Table Detection), a manually curated dataset of 2,323 table-containing pages spanning 18 languages and 15 scripts across diverse domains. We also propose EMBLEM, a masking-based paradigm for Multi-script Table Detection (MTD). EMBLEM generates masked images that conceal script- and font-specific details, enabling models pre-trained on abundant English documents to focus on script-agnostic page layout. Experiments across three table detection architectures show that EMBLEM consistently outperforms strong baselines on MANDALA while remaining competitive on five standard English-dominant benchmarks. Using only English masked images for fine-tuning, with no multi-script training data, EMBLEM achieves an absolute F1-score gain of 20.8% on MANDALA. We release MANDALA along with the accompanying code and models at https://github.com/IITB-LEAP-OCR/EMBLEM.git.

CommentsAccepted in International Conference on Document Analysis and Recognition 2026

DOI:10.1007/978-3-032-36033-5_3

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑