发表机构
School of Informatics; University of Edinburgh(信息学院; 爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对多表格文档问答提出无需训练的两步像素级表格压缩方法,在长文档上节省41%总Token,比高效单步压缩配置少用15%Token且无准确率损失,还较原生分辨率单步问答提升7个百分点准确率。
AI 中文摘要
针对真实世界文档的问答需要处理文本与表格交织的长输入。将上下文表示为图像的光学上下文压缩有望降低Token成本,但它对表格理解的影响尚不明确。本文研究面向多表格文档问答的像素级表格压缩,在两个基准和五个视觉Token预算下评估了五个视觉语言模型(VLM)。将表格以原生分辨率表示为图像,在性能和效率上都与文本相当;但缩小表格会使模型用更长、效果更差的推理路径来补偿可读性损失,抵消了预期的节省。不过,高度缩小的表格仍保留了足够的信号,以识别其是否与问题相关。我们利用这种不对称性,提出一种无需训练的两步方法:模型首先从像素压缩的上下文中识别回答问题所需的表格,然后以原生分辨率对这些表格进行推理。在长文档上,该方法节省了41%的总Token,与使用原生分辨率表格的单步问答相比,准确率提升了7个百分点;同时,它比最有效的单步压缩配置少用15%的Token,且无准确率损失。
英文摘要
Answering questions over real-world documents requires processing long inputs that interleave text with tables. Optical context compression, which represents context as images, promises to reduce token cost, but its effect on table understanding remains unclear. We study pixel-level table compression for question answering over documents with multiple tables, evaluating five VLMs across two benchmarks and five visual-token budgets. Representing tables as images at native resolution matches text in both performance and efficiency, but downscaling them makes models compensate the loss in readability with longer, less effective reasoning traces that cancel the expected savings. Highly downscaled tables, however, preserve enough signal to identify whether they are relevant to a question. We exploit this asymmetry with a training-free, two-step method: the model first identifies the tables needed to answer a question from a pixel-compressed context, and then reasons over those at native resolution. On long documents, our method saves 41% of total tokens and gains 7 accuracy points over single-step QA with native resolution tables. It also uses 15% fewer tokens than the most efficient single-step compressed configuration, with no accuracy loss.