超越准确率:模板化文档提取中视觉语言模型的鲁棒性、成本与治理权衡
Beyond Accuracy: Robustness, Cost, and Governance Trade-offs for Vision-Language Models in Templated Document Extraction
浏览论文内容
中文总结 AI 辅助
本研究通过测量11个视觉语言模型在模板化文档提取任务上的表现,提出一个基于任务概况(质量、延迟、治理、数量)的选择框架,指导实践者权衡鲁棒性、成本与治理,并开源发布相关资源。
中文摘要 AI 辅助
视觉语言模型(VLMs)越来越多地被用于从商业文档中提取结构化字段,然而大多数评估仅报告干净基准上的准确率,对于实践者根据给定任务复杂度选择方法提供的指导甚少。我们通过一项基于测量的研究和开源发布来填补这一空白。在由750份合成支票文档组成的保留测试集上,对十一个系统(包括三个商业系统、两个推理系统、五个预训练和微调形式的开源VLM,以及一个非LLM的OCR到正则表达式基线)进行评分,基于3K样本的微调使最佳开源VLM的F1分数提升至0.98以上,在此任务上超过了所有零样本商业系统,而GPT-5在F1上领先商业系统,Claude Sonnet 4.5在日期字段上表现崩溃。为了将这些测量结果转化为可操作的决策,我们引入了一个面向实践者的选择框架,该框架通过过滤和总成本最小化,将任务概况(质量、延迟、治理、数量)映射到推荐方法,并以一个假设的中等规模文档提取场景为例进行说明。
英文摘要
Vision-language models (VLMs) are increasingly used to extract structured fields from business documents, yet most evaluations report accuracy on clean benchmarks and offer little guidance to practitioners choosing an approach for a given task complexity. We address this gap with a measurement-grounded study and an open-source release. Across eleven systems (three commercial, two reasoning, five open-source VLMs in pretrained and fine-tuned form, and a non-LLM OCR->regex floor) scored on a 750-document held-out pool of synthetic checks, fine-tuning on 3K samples lifts the best open-source VLMs above F1 0.98-above every zero-shot commercial system on this task-while GPT-5 leads the commercial pool on F1 and Claude Sonnet 4.5 collapses on Date. To turn these measurements into actionable choices, we introduce a practitioner-oriented selection framework that maps a task profile (quality, latency, governance, volume) to a recommended approach via filtering and total-cost minimization, illustrated on a hypothetical mid-volume document-extraction scenario.
发表机构
- John Hancock(约翰汉考克)
机构由 AI 辅助整理,请以论文原文为准。