arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

表格解码:DELTA 用于结构,TARQA 用于理解

Tables Decoded: DELTA for Structure, TARQA for Understanding

Jahanvi Rajput, Dhruv Kudale, Saikiran Kasturi, Utkarsh Verma, Ganesh Ramakrishnan

arXiv 2609.17458首次发表:更新:

发表机构

Indian Institute of Technology Bombay; BharatGen(印度理工学院孟买分校; BharatGen)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对表格理解,提出基于结构化文本的DELTA和TARQA,分别处理结构识别和问答,在多个基准上达到先进性能,并验证了多语言鲁棒性。

AI 中文摘要

表格理解是文档智能中的核心任务,涵盖两个关键子任务:表格重建和表格视觉问答(TabVQA)。虽然近期方法主要依赖基于表格图像的视觉-语言模型(VLM),我们提出了一种基于结构化文本表示的更具可扩展性和有效性的替代方案。这些表示更易于处理,与大型语言模型(LLM)更自然地对齐,并消除了对特定语言视觉编码器的需求,使其特别适用于多语言文档。我们提出了DELTA,它将物理结构识别、逻辑结构识别和OCR分离,以准确提取布局和内容。DELTA以优化表格结构语言(OTSL)输出表格,这是一种紧凑且统一的格式,用于编码单元格排列和文本内容。在表格结构识别(TSR)上,DELTA在FinTabNet、PubTabNet和PubTables-1M上取得了与最先进方法相当的TEDS-Structure分数。我们进一步通过我们策划的印地语基准TORQUE证明了其在非英语表格上的鲁棒性。在此基础上,我们引入了TARQA,一个在OTSL序列上微调的LLM。我们的方法在WTQ(TabQA)上取得了9.3个百分点的提升,在FinTabNetQA(TabVQA)上取得了9.2个百分点的提升。在TORQUE上,我们的方法在所有VLM和DELTA + LLM变体中排名第二。我们发布了我们的代码、模型和基准,网址为:此https URL

英文摘要

Table understanding is a core task in document intelligence, encompassing two key subtasks: table reconstruction and table visual question answering (TabVQA). While recent approaches predominantly rely on vision- language models (VLMs) operating on table images, we propose a more scalable and effective alternative based on structured textual representations. These representations are easier to process, align more naturally with LLMs, and eliminate the need for language-specific visual encoders, making them particularly suitable for multilingual documents. We present DELTA, which separates physical structure recognition, logical structure recognition, and OCR to extract both layout and content accurately. DELTA outputs tables in Optimised Table Structure Language (OTSL), a compact and unified format that encodes cell arrangements and textual content. On table structure recognition (TSR), DELTA achieves TEDS- Structure scores comparable with state-of-the-art methods across FinTabNet, PubTabNet, and PubTables-1M. We further establish its robustness on non-English tables through our curated Hindi benchmark, TORQUE. Building on this, we introduce TARQA, an LLM fine-tuned on OTSL sequences. Our approach yields gains of 9.3 p.p. on WTQ (TabQA) and 9.2 p.p. on FinTabNetQA (TabVQA), respectively. On TORQUE, our method ranks second among all VLMs and DELTA + LLM variants. We release our code, models, and benchmark at: https://github.com/Tihiitborg/Tables-Decoded

CommentsAccepted at the IEEE/CVF Winter Conference on Applications of Computer Vision 2026

DOI:10.1109/WACV61042.2026.00272

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑