在高风险公共部门应用中使用开源模型评估结构化信息提取
Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application
浏览论文内容
中文总结 AI 辅助
本研究针对欧盟AI法案高风险公共部门应用场景,构建基准评估开源OCR、LLM、VLM在学生申请处理任务的性能,发现零样本下多数模型表现差,输入结构保留是关键因素。
中文摘要 AI 辅助
从非结构化文档中提取结构化信息是各行业数字化转型的关键组成部分。尽管专有解决方案在商业应用中占据主导地位,但不断发展的开源光学字符识别(OCR)引擎、大语言模型(LLM)以及视觉语言模型(VLM)生态系统提供了可及的替代方案。然而,针对现实多步骤提取流程的系统评估仍较为匮乏。负责任地使用此类提取工具需要针对现实任务进行全面评估,尤其是这些解决方案将成为欧盟《人工智能法案》归类为高风险的公共部门应用的关键组件。为填补这一空白,我们提出了一个全面基准,评估开源系统在被归类为高风险的复杂现实文档处理任务——国际学习项目的学生申请——上的端到端性能。我们对最先进的OCR引擎、LLM和VLM进行了全面实证评估。结果显示,尽管VLM总体上优于OCR+LLM流程,但即使是最先进的开源模型在零样本设置下也难以可靠地处理此类任务。35种配置中仅有4种的F1分数超过0.5,最佳OCR+LLM流程与顶级VLM性能相当,不过大多数OCR+LLM组合的表现要差得多。约75%的配置得分低于0.25。模型规模会影响性能,但这种关系是非线性的:更大规模的模型并不能保证成比例的更好结果。输入质量,尤其是OCR输出的结构保留,是独立于下游模型能力的关键因素。
英文摘要
The extraction of structured information from unstructured documents represents a critical component of digital transformations in all sectors. While proprietary solutions dominate commercial applications, a rapidly growing ecosystem of open-source Optical Character Recognition (OCR) engines, Large Language Models (LLMs), and Vision-Language Models (VLMs) offers accessible alternatives. However, systematic evaluations on realistic, multi-step extraction pipelines remain scarce. Responsible usage of such extraction tools require comprehensive evaluations on realistic tasks, especially as these solutions will be key components of applications in the public sector that the EU AI act categorizes as high risk. To address this gap we present a comprehensive benchmark assessing the end-to-end performance of open-source systems on a complex real-world document processing task classified as high risk: Student applications for an international study program. We conduct a comprehensive empirical evaluation with state-of-the-art OCR engines, LLMs and VLMs. Our results reveal that while VLMs generally outperform OCR+LLM pipelines, even state-of-the-art open-source models struggle to handle such tasks reliably in zero-shot settings. Only 4 of 35 configurations achieved F1 scores above 0.5, with the best OCR+LLM pipeline matching top VLM performance, though most OCR+LLM combinations performed substantially worse. Roughly 75\% of all configurations scored below 0.25. Model scale influences performance, yet the relationship is non-linear: substantially larger models do not guarantee proportionally better results. Input quality, particularly the structural preservation of OCR output, emerges as a critical factor independent of downstream model capability.
发表机构
- Berlin University of Applied Sciences(柏林应用科学大学)
- Einstein Center Digital Future(爱因斯坦数字未来中心)
机构由 AI 辅助整理,请以论文原文为准。