AI 中文总结
研究针对企业文档模式引导抽取的评估需求,构建首个同时评估值准确率等多指标的基准ExtractBench,发现LlamaExtract Agentic Plus在成本远低于编码智能体的情况下,准确率可与之媲美。
AI 中文摘要
企业工作流越来越依赖智能体进行模式引导的抽取:给定一份文档和用户定义的模式,智能体会严格遵循该模式生成正确输出,并附带作为基础元数据的源证据。我们提出ExtractBench,这是一个用于模式引导抽取的基准,据我们所知,它是首个同时对值准确率、大规模记录完整性、基础信息和测量成本进行评分的基准。评估系统包含370份企业文档共4869页,覆盖8个业务领域和67种文档类型,带有清晰标签以区分其挑战场景。可扩展的模式和真值整理流程结合了真实文档的独立系统一致性、合成列表的已知值以及表单的人工验证。我们报告了用于值准确率的不依赖顺序的值F1,以及两个用于源可追溯性的基础指标:词级和页级F1。商业视觉语言模型(VLM)在短文档上表现良好,但在长文档上常截断记录列表;而编码智能体在更高成本下保持更高准确率。LlamaExtract Agentic Plus在所有三个指标中排名第一,其准确率可与编码智能体媲美,但成本仅为其一小部分。数据集和评估代码可在HuggingFace和GitHub上获取。
英文摘要
Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and measured cost together. The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating their challenge scenarios. The scalable schema and ground-truth curation pipeline combines independent-system agreement for real documents, known values for synthetic lists, and human verification for forms. We report order-insensitive value F1 for value accuracy, plus two grounding metrics for source traceability: word- and page-level F1. Commercial VLMs perform well on short documents but often truncate record lists on long ones, while coding agents retain higher accuracy at much higher cost. LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost. Dataset and evaluation code are available on \href{https://huggingface.co/datasets/llamaindex/ExtractBench}{HuggingFace} and \href{https://github.com/run-llama/ExtractBench}{GitHub}.