arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20868cs.CVcs.CL

识别、定位、关联:从文档图像中进行端到端键值提取

Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images

A. Said Gurbuz, Ahmed Nassar, Christoph Auer, Maksym Lysak, Lucas Morin, Matteo Omenetti, Tim Strohmeyer, Panagiotis Vagenas, Nikolaos Livathinos, Michele Dolfi, Peter Staar

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对传统文档处理的误差传播问题,微调SmolDocling模型实现端到端键值提取,在多个数据集上性能优于更大基线模型,且体积小、推理快。

中文摘要 AI 辅助

传统文档处理流程将光学字符识别(OCR)引擎与下游结构化信息提取模型级联,导致多阶段误差传播。我们对SmolDocling(一个拥有2.56亿参数的紧凑视觉语言模型VLM)进行微调,实现从文档图像直接进行端到端键值提取,单次完成识别、定位与关联,无需OCR预处理。我们对DocTags进行扩展,新增专门的键、值、区域及关联标签,使统一输出序列支持多对多关系。为解决数据限制,我们设计了增强流水线,结合合成表单填充与保留完整键值子图的基于图的裁剪操作。我们进一步引入布局感知评估框架,将文本匹配扩展至空间边界框验证。在FUNSD、XFUND及一个大规模私有数据集上,我们的模型在布局感知评估下优于更大的零样本VLM基线,同时比Qwen2.5-VL(7B)小27倍,推理速度快5倍以上。模型权重将在发表后公开发布。

英文摘要

Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information extraction, leading to multi-stage error propagation. We fine-tune SmolDocling, a compact 256M-parameter vision-language model (VLM), to perform end-to-end key-value extraction directly from document images, jointly solving identification, localization, and association in a single pass without OCR preprocessing. We extend DocTags with specialized key, value, region, and link tags, enabling many-to-many relationships in a unified output sequence. To address data limitations, we design an augmentation pipeline combining synthetic form filling and graph-based crops that preserve complete key-value subgraphs. We further introduce a layout-aware evaluation framework extending text matching with spatial bounding box verification. On FUNSD, XFUND, and a large-scale private dataset, our model outperforms larger zero-shot VLM baselines under layout-aware evaluation, while being 27 times smaller than Qwen2.5-VL (7B) and over 5 times faster at inference. The model weights will be released publicly after publication.

发表机构

  • IBM Research Zurich(IBM苏黎世研究院)
  • ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑