发表机构
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有文档到表格基准无法适配关系数据库构建需求的问题,推出Doc2DB-Bench基准,含多领域长文档实例,为基于LLM的可靠数据系统提供测试平台。
AI 中文摘要
实用AI系统日益需要将长篇、异构文档转化为可查询的关系数据库,而非孤立的电子表格。在金融、医疗、教育、交通及企业运营等领域,下游工作流依赖规范化模式、实体标识、键、跨表关系与完整性约束,以开展分析、合规、审计及基于SQL的决策。现有的文档到表格基准不足以适配该场景:将证据扁平化至单张表格会导致实体重复、掩盖多对多关系、生成稀疏记录,且无法检验提取的事实是否构成有效数据库实例。这催生了对文档理解任务的需求,即应将其视为数据库构建而非字段提取。我们推出Doc2DB-Bench,这是一个文档到数据库构建的基准,包含42种模式、7个领域组的203个长文档实例,含117个实体表、132个关系表、7341行及41935个单元格。该基准通过可控的DB-to-Doc合成流水线构建,按表内提取与表间推理的分类法组织,生成的文档经真实性验证,证明与现实参考资料无法区分。因此,Doc2DB-Bench为可靠、可审计且关系忠实的基于LLM的数据系统提供了测试平台,该基准可通过此URL公开获取。
英文摘要
Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domains such as finance, healthcare, education, transportation, and enterprise operations, downstream workflows rely on normalized schemas, entity identities, keys, cross-table relationships, and integrity constraints for analytics, compliance, auditing, and SQL-backed decision making. Existing Document-to-Table benchmarks are insufficient for this setting: flattening evidence into single tables can duplicate entities, obscure many-to-many relationships, create sparse records, and avoid testing whether extracted facts form a valid database instance. This creates an urgent need to evaluate document understanding as database construction rather than field extraction. We introduce Doc2DB-Bench, a benchmark for Document-to-Database construction, containing 203 long-document instances across 42 schemas and seven domain groups, with 117 entity tables, 132 relationship tables, 7,341 rows, and 41,935 cells. Built through a controllable DB-to-Doc synthesis pipeline and organized by a taxonomy of intra-table extraction and inter-table reasoning, the generated documents undergo authenticity verification, proving indistinguishable from real-world references. Doc2DB-Bench thus provides a testbed for reliable, auditable, and relationally faithful LLM-based data systems. The benchmark is publicly available at https://github.com/SetonLiang/Doc2DB-Bench.
Comments24 pages, 13 figures, 7 tables