arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16050cs.IRcs.DBcs.LG

面向可扩展电子表格表格理解的结构化预测:从单元格类型到表格范围(扩展版)

Structured Prediction for Scalable Spreadsheet Table Understanding: From Cell Types to Table Ranges (Extended Version)

Antoine Gauquier, Ioana Manolescu, Pierre Senellart

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对电子表格理解的单元格类型分类与表格检测任务,提出结合CRF-LightGBM的两阶段流水线,在公开基准StatSheets上实现了与先进模型相当的性能,且计算资源需求更低。

中文摘要 AI 辅助

电子表格是发布表格数据的主要媒介,但由于布局异构、文件格式多样及组织规范不一致,自动从中提取结构化内容仍存在困难。我们解决电子表格理解中的两个核心任务:单元格类型分类(Cell-Type Classification, CTC),即给单元格分配角色;表格检测(Table Detection, TD),即识别工作表内的表格边界框。我们提出一种高效的两阶段流水线,其中学习得到的CTC模型为确定性TD算法提供输入。对于CTC,我们采用基于65个结构化特征的LightGBM分类器,搭配成对条件随机场(pairwise CRF)以强制单元格网格的空间一致性。我们的TD方法通过确定性五阶段流程从预测的单元格类型中提取表格范围。为评估,我们构建并公开了StatSheets,这是一个多语言基准,包含来自14个公共数据提供商、覆盖多个国家及多种文件格式的737份手动标注工作表。在5折交叉验证下,我们的CRF-LightGBM系统在CTC任务上达到平均文件宏F1值0.937,与基于GPU的TUTA Transformer仅相差0.6个百分点,同时所需计算资源显著更少。对于TD任务,我们的确定性方法优于基于区域的基线,且与近期基于大语言模型(LLM)的系统如SpreadsheetLLM相比仍具竞争力。这些结果表明,将非线性结构化预测与确定性范围提取相结合,可提供一种具有竞争力、可扩展且计算高效的电子表格表格理解方法。

英文摘要

Spreadsheets are a primary medium for publishing tabular data, yet automatically extracting structured content from them remains difficult due to heterogeneous layouts, diverse file formats, and inconsistent organizational conventions. We address two core tasks in spreadsheet understanding: Cell-Type Classification (CTC), which assigns roles to cells, and Table Detection (TD), which identifies table bounding boxes within sheets. We propose an efficient two-stage pipeline in which a learned CTC model feeds a deterministic TD algorithm. For CTC, we use a LightGBM classifier over 65 structured features together with a pairwise CRF enforcing spatial consistency across the cell grid. Our TD method extracts table ranges from predicted cell types by a deterministic five-stage procedure. For evaluation, we built and share StatSheets, a multilingual benchmark of 737 manually annotated sheets from 14 public data providers across multiple countries and file formats. Under 5-fold cross-validation, our CRF-LightGBM system achieves a Mean File-Macro F1 score of 0.937 on CTC, within 0.6 percentage points of the GPU-based TUTA Transformer, while requiring substantially fewer computational resources. For TD, our deterministic approach outperforms region-based baselines and remains competitive with recent LLM-based systems such as SpreadsheetLLM. These results demonstrate that combining non-linear structured prediction with deterministic range extraction provides a competitive, scalable, and computationally efficient approach to spreadsheet table understanding.

发表机构

  • DI ENS, ENS, CNRS, PSL, Inria(法国高等师范学院(DI ENS)、巴黎高等师范学院、法国国家科学研究中心、PSL大学、法国国家信息与自动化研究所)
  • Inria & Institut Polytechnique de Paris(法国国家信息与自动化研究所、巴黎综合理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑