XL-DocBench:基于证据的超长文档理解基准测试
XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding
- Wuhan University(武汉大学)
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
XL-DocBench是含1519个问题、最长2303页的人工验证超长文档理解基准,覆盖多页证据、结构化内容等场景,可用于评估LLMs在专业文档理解中的实际表现。
AI中文摘要:
现实世界的文档任务常要求专业人员回答来自年度报告、法规、临床指南和技术手册等数百或数千页内容的问题,部分问题还需对比相关报告。因此,可靠的长文档理解是将大语言模型(LLMs)应用于合规、临床、金融和工程工作流的前提——此类工作流中,决策必须可追溯至具体证据页,无依据回答的成本极高;但多数现有基准仍仅针对短上下文或单页问答。我们推出XL-DocBench,这是一个经人工全面验证的超长文档理解基准,涵盖6个专业领域的1519个保留问题,上下文最长达2303页。XL-DocBench超越了单页查找:1103个示例(占72.6%)需使用多页证据,最终集合还包含556个涉及表格、图表或图形的问题(占36.6%),以及165个需来自多份文档证据的问题(占10.9%)。每个问题都有12种推理标签之一、专家标注的证据页、类型化验证规则和答案格式,其中包含218个无答案(None)案例。我们通过树引导合成流程构建该基准,再经人工制品筛选及194位专家的全面验证。通过结合超长专业上下文、单页证据与类型化规则,XL-DocBench填补了现有单页、短多页或纯文本长上下文基准的空白,使后续研究能将系统失败归因于检索、证据使用或规则遵循,而非单一排行榜分数。结果显示,当前系统在处理长上下文、多页证据及专业文档的结构化推理方面仍存在困难。
英文摘要:
Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands of pages. Some questions also require comparing related reports. Reliable long-document understanding is therefore a prerequisite for using LLMs in compliance, clinical, financial, and engineering workflows, where decisions must be traceable to specific evidence pages and the cost of an unsupported answer is high -- yet most existing benchmarks still measure short-context or single-page QA. We introduce XL-DocBench, a fully human-verified benchmark for extra-long document understanding, with 1,519 retained questions from six professional domains and contexts up to 2,303 pages. XL-DocBench goes beyond page-level lookup. 1,103 examples (72.6\%) use multiple evidence pages. The final set also includes 556 questions (36.6\%) that use tables, charts, or figures, and 165 questions (10.9\%) that require evidence from multiple documents. Each question has one of twelve reasoning labels, expert-annotated evidence pages, a typed verification rule, and an answer format, including 218 None-answer cases. We build the benchmark with a tree-guided synthesis pipeline followed by artifact filters and full verification by 194 human experts. By coupling extra-long professional contexts with page-level evidence and typed rules, XL-DocBench fills a gap left by prior single-page, short multi-page, or text-only long-context benchmarks, and lets future work attribute system failures to retrieval, evidence use, or rule following rather than to a single leaderboard score. The results show that current systems still struggle with long contexts, multi-page evidence, and structured reasoning over professional documents.