识别、定位、关联:从文档图像中进行端到端键值提取
Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images
浏览论文内容
中文总结 AI 辅助
该研究针对传统文档处理的误差传播问题,微调SmolDocling模型实现端到端键值提取,在多个数据集上性能优于更大基线模型,且体积小、推理快。
中文摘要 AI 辅助
传统文档处理流程将光学字符识别(OCR)引擎与下游结构化信息提取模型级联,导致多阶段误差传播。我们对SmolDocling(一个拥有2.56亿参数的紧凑视觉语言模型VLM)进行微调,实现从文档图像直接进行端到端键值提取,单次完成识别、定位与关联,无需OCR预处理。我们对DocTags进行扩展,新增专门的键、值、区域及关联标签,使统一输出序列支持多对多关系。为解决数据限制,我们设计了增强流水线,结合合成表单填充与保留完整键值子图的基于图的裁剪操作。我们进一步引入布局感知评估框架,将文本匹配扩展至空间边界框验证。在FUNSD、XFUND及一个大规模私有数据集上,我们的模型在布局感知评估下优于更大的零样本VLM基线,同时比Qwen2.5-VL(7B)小27倍,推理速度快5倍以上。模型权重将在发表后公开发布。
英文摘要
Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information extraction, leading to multi-stage error propagation. We fine-tune SmolDocling, a compact 256M-parameter vision-language model (VLM), to perform end-to-end key-value extraction directly from document images, jointly solving identification, localization, and association in a single pass without OCR preprocessing. We extend DocTags with specialized key, value, region, and link tags, enabling many-to-many relationships in a unified output sequence. To address data limitations, we design an augmentation pipeline combining synthetic form filling and graph-based crops that preserve complete key-value subgraphs. We further introduce a layout-aware evaluation framework extending text matching with spatial bounding box verification. On FUNSD, XFUND, and a large-scale private dataset, our model outperforms larger zero-shot VLM baselines under layout-aware evaluation, while being 27 times smaller than Qwen2.5-VL (7B) and over 5 times faster at inference. The model weights will be released publicly after publication.
发表机构
- IBM Research Zurich(IBM苏黎世研究院)
- ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。