arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32660cs.CVcs.AI

InterTab:面向多模态表格推理的交错视觉-结构对齐

InterTab: Interleaved Visual-Structure Alignment for Multi-Modal Table Reasoning

Hanqian Li, Sirui Huang, Chen Ling, Jungang Li, Yu Huang, Kening Zheng, Yonghua Hei, Xiangrong He, Shiyi Wang, Pengcheng Zhu, Dongnan Liu, Wei Zhou, Linjian Mo, Nai Ding, Xuming Hu

首次发表
浏览论文内容

中文总结 AI 辅助

InterTab提出交错结构感知框架,通过工具调用裁剪表格区域,结合两阶段训练,在九个基准上将平均准确率提升至73.17%。

中文摘要 AI 辅助

表格图像保留了在文本序列化中常丢失的结构信息,对其进行推理需要逐步定位相关的行、列和单元格。当前的多模态大语言模型(MLLMs)在推理前对整个图像进行一次编码,因此无法随着问题的展开获取行、列和单元格级别的证据。编码器侧的表格结构和通用的交错视觉思维链仍未能将每个推理步骤绑定到该结构上。我们提出了InterTab,一种用于表格图像上思维链推理的交错结构感知框架,它将思维链与裁剪结构对齐表格区域的工具调用交错进行。首先,我们构建了InterTab-22K,其中包含推理轨迹,每一步都关联到一个结构位置和一个边界框。InterTab分两个阶段训练:在InterTab-22K上进行监督结构感知对齐(SSA),教会模型将推理与结构对齐的裁剪交错进行;主动定位优化(ALO)进一步优化答案正确性、定位IoU和输出格式,同时惩罚缺失或多余的工具调用。在九个表格基准上的实验表明,InterTab将其骨干模型的平均准确率从68.28%提升到73.17%,并在所有比较方法中取得了最佳平均性能。代码和数据将很快发布。

英文摘要

Table images preserve structural information that are often lost in text serialization, and reasoning over them requires locating relevant rows, columns, and cells step by step. Current multimodal large language models (MLLMs) encode the whole image once before reasoning, so they cannot pick up row-, column-, and cell-level evidence as the question unfolds. Encoder-side table structure and generic interleaved visual chain-of-thought still do not bind each reasoning step to that structure. We propose \textbf{InterTab}, an \textbf{Inter}leaved structure-aware framework for CoT reasoning over \textbf{Tab}le images, interleaves chain-of-thought with tool calls that crop structure-aligned table regions. First, we build InterTab-22K, includes reasoning trajectories in which each step is tied to both a structural location and a bounding box. InterTab is trained in two stages: supervised structure-aware alignment (SSA) on InterTab-22K teaches the model to interleave reasoning with structure-aligned crops, and active localization optimization (ALO) further optimizes answer correctness, localization IoU, and output format, while penalizing missing or excessive tool calls. Experiments on nine table benchmarks show that InterTab improves the average accuracy of its backbone from 68.28% to 73.17% and achieves the best average performance among all compared methods. Code and data will be released soon.

↑