发表机构
Citigroup Inc.; Ernst & Young LLP; NVIDIA Corporation(花旗集团; 安永会计师事务所; 英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一个生产就绪的内容提取系统,通过选择性OCR路由、稀缺优先策展、结构感知分块和检索评估器,使生成式AI的数据摄取阶段显式化、可配置且可衡量,在180篇文档上达到97.4分提取得分和68.6%的hit@1。
AI 中文摘要
机器学习的演进已逐步改变了智能在AI系统中的驻留位置。在传统机器学习中,任务、数据表示、标签和模型架构紧密耦合,因此数据准备是狭窄的、受模式约束且可见的。生成式AI将模型与任何单一任务解耦:一个基础模型服务于开放式的下游任务,而模型侧获得的通用性以数据侧的异构性为代价,因为企业知识以人们使用的格式(PDF、演示文稿、电子表格、扫描文档、表单、表格、图表和混合版式文件)编写,这些格式同时携带文本、视觉、几何和结构信息。语言模型或检索器无法在此接口上对错误表示的信息进行可靠推理,因此内容提取成为其自身的生命周期阶段,其错误是任何下游检索器或重排序器都无法修复的。本文提出了一个生产就绪的内容提取系统,使该阶段显式化、可配置且可衡量。该系统包含选择性OCR路由、稀缺优先的策展引擎(带有基于引用的提取评分器,用于衡量字符、单词和表格结构准确率)、确定性的结构感知父子分块器,以及只读检索评估器(从每一页生成基于事实的问题,并报告Hit@k、平均倒数排名和延迟)。在180篇文档的语料库上,最佳提取器得分为97.4分(满分100分)(字符错误率0.13%,表格相似度0.995),分块器在25,050个生成问题上达到hit@1为68.6%、hit@10为92.8%、MRR为0.77。我们提炼出三条设计原则(结构先于语义、绝不修改所测量之物、预算你的标签),并将可衡量的内容提取定位为企业智能体系统的感知层。
英文摘要
The evolution of machine learning has progressively changed where intelligence resides in an AI system. In conventional machine learning the task, data representation, labels and model architecture were tightly coupled, so data preparation was narrow, schema-bound and visible. Generative AI decouples the model from any single task: one foundation model serves open-ended downstream tasks, and the generality gained on the model side is matched by heterogeneity on the data side, because enterprise knowledge is authored in the formats people use (PDF, presentations, spreadsheets, scanned documents, forms, tables, diagrams and mixed-layout files) that carry textual, visual, geometric and structural information at once. A language model or retriever cannot reason reliably over information misrepresented at this interface, so content extraction becomes a lifecycle stage in its own right whose errors no downstream retriever or re-ranker can repair. This paper presents a production-ready content-extraction system that makes this stage explicit, configurable, and measurable. The system incorporates selective OCR routing, a scarcity-first curation engine with a reference-based extraction scorer that measures character, word, and table-structure accuracy, a deterministic structure-aware parent-child chunker, and a read-only retrieval evaluator that generates grounded questions from every page and reports Hit@k, mean reciprocal rank, and latency. On a 180-document corpus the best extractor scores 97.4 of 100 (character error rate 0.13%, table similarity 0.995) and the chunker reaches hit@1 of 68.6%, hit@10 of 92.8% and MRR 0.77 over 25,050 generated questions. We distil three design principles (structure before semantics, never mutate what you measure, budget your labels) and position measured content extraction as the perception layer of enterprise agentic systems.