arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过边界框引导增强表格结构识别

Enhancing Table Structure Recognition via Bounding Box Guidance

Lei Hu, Shuangping Huang

arXiv 2609.08705首次发表:更新:

发表机构

South China University of Technology(华南理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对图像到序列方法忽略边界框信息导致复杂场景错误的问题,提出BGTR框架,先预测单元格边界框再引导HTML生成,并引入合成数据集SNSTab和渐进训练,在五个基准上取得SOTA性能。

AI 中文摘要

表格结构识别(TSR)旨在从表格图像中提取单元格的边界框和表格结构(例如HTML)。尽管当前方法已取得显著进展,但最新的图像到序列方法在预测HTML序列时忽略了边界框信息的显式利用,导致在复杂场景中出现错误预测。在本文中,我们提出了一种新颖的框架BGTR(边界框引导的表格识别器)。为了更有效地利用边界框信息,我们首先预测单元格的边界框,然后使用这些信息来引导HTML序列的生成。虽然利用边界框信息可以提高HTML序列的准确性,但对于自然场景表格,数据量过小,无法充分训练边界框引导的HTML生成。为此,我们对自然场景表格采用渐进式训练方法,并引入了SNSTab,一个合成生成的自然场景表格数据集。我们在五个基准数据集上的实验证明了最先进的性能。

英文摘要

Table Structure Recognition (TSR) aims to extract the bounding boxes of cells and table structure (e.g., HTML) from table images. Although current approaches have made significant progress, the latest image-to-sequence methods overlook the explicit utilization of the bounding box information when predicting HTML sequences, leading to error predictions in complex scenes. In this paper, we introduce a novel framework BGTR (Bounding Box-Guided Table Recognizer). To more effectively utilize bounding box information, we first predict the bounding boxes of cells and then use this information to guide the generation of HTML sequences. While utilizing bounding box information can enhance the accuracy of HTML sequences, for natural scene tables, the data volume is too small to allow for sufficient training of bbox-guided HTML generation. In response, we adopt a progressive training method for natural scene tables and introduce SNSTab, a synthetically generated natural scene table dataset. Our experiments on five benchmark datasets demonstrate SOTA performance.

CommentsICPR 2024. Upload for archiving

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑