从像素到配对:噪声文档环境下基于LLM的键值提取综合基准
From Pixels to Pairs: A Comprehensive Benchmark of LLM-Driven Key-Value Extraction in Noisy Document Settings
- University of Texas at Arlington(德克萨斯大学阿灵顿分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出一个基准,系统评估LLM在噪声OCR文档中的键值提取,发现语义推理与文本保真度是关键,且噪声下模型差距缩小。
AI中文摘要:
大型语言模型(LLMs)越来越多地被用于从文档中进行结构化信息提取,但它们在真实OCR噪声下的行为仍鲜为人知。我们提出了一个系统性的基准,用于评估开源指令微调LLM在干净文本和噪声OCR条件下的键值对(KVP)提取性能。我们在FUNSD、CORD和SROIE基准上,使用Gold文本标注以及来自PaddleOCR、EasyOCR和Tesseract的OCR输出,评估了具有代表性的仅解码器模型(Gemma、Mistral、Qwen2.5、LLaMA 3和DeepSeek)。一个统一的评估协议在一致条件下隔离了输入质量、模型设计和提示的影响。结果表明,当高质量文本可用时,现代LLM作为强大的语义提取器,在某些情况下接近监督的布局感知系统。然而,在OCR噪声下,性能显著下降,并且随着输入损坏的增加,模型之间的性能差距缩小。在所有数据集中,提取性能由两个因素决定:对文本的语义推理以及在OCR噪声下保持文本保真度。虽然较大的模型在干净文本上改善了结果,但这些收益在噪声输入下减弱,此时OCR质量成为主导因素。我们还识别了反复出现的失败模式,包括键值错位、幻觉和数字损坏。我们的发现凸显了干净文本评估与真实世界部署之间的差距,强调了联合改进OCR质量、结构推理和基于LLM的语义建模的必要性。
英文摘要:
Large language models (LLMs) have demonstrated strong capabilities in document key-value pair (KVP) extraction, yet controlled evaluations of their robustness to optical character recognition (OCR) output remain limited. This leaves an important gap in understanding their reliability in real-world OCR-to-LLM pipelines. Unlike end-to-end Vision-Language Models (VLMs), which jointly perform visual perception and semantic extraction, modular pipelines allow these stages and their errors to be isolated and audited. We introduce a controlled benchmark that distinguishes downstream LLM extraction behavior from upstream OCR degradation. It evaluates 136 experimental configurations and 17,688 document-level inferences generated with deterministic decoding across five instruction-tuned open-weight LLMs (2B-8B parameters), three datasets, and four text-quality conditions. The evaluation combines a full zero-shot comparison, targeted one- to three-shot experiments, and a sensitivity analysis of 40 configurations across 20 frozen demonstration sets. By separating Key Recall (annotated-field recovery) from Exact Match and Value F1 (exact and partial value recovery, respectively), we test whether OCR degradation affects field identification and value reproduction differently across models. Our findings challenge three practical assumptions: (1) clean-text performance reliably predicts real-world robustness, (2) model rankings remain consistent across annotation-derived Gold and OCR-derived text, and (3) additional few-shot demonstrations monotonically improve extraction accuracy. The observed model-ranking reversals and unstable few-shot gains expose important reliability risks under noisy document conditions. We release the benchmarking framework, dataset splits, and evaluation scripts to support reproducible research.