arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为基于大语言模型分析的长文档建立索引

Indexing Long Documents for LLM-Based Analysis

Donna Pham

arXiv 2608.21237首次发表:更新:

AI 中文总结

针对LLM分析长文档时速度慢、成本高、易幻觉等问题,提出一种由LLM自动构建的分层纯文本索引,在NarrativeQA数据集上达到与DocETL相当准确率且成本降低40%。

AI 中文摘要

临床记录、法律合同、科学论文等长文档正越来越多地使用大语言模型(LLM)进行分析。显然,将完整文档输入模型以处理每个问题,最终会变得缓慢、成本高昂、易产生幻觉,且无法在不同问题间复用工作。我们探索了一种基于索引的文档分析解决方案,提出了一种分层纯文本索引,该索引为每个文档构建一次,后续查询可调用它。受经典B+树启发,该索引将文档组织成页面,从根节点的通用摘要到叶节点的具体摘要排列,但它在三个方面不同于B+树,以适配LLM访问:内容存在于每个层级、索引是纯文本而非属性值、导航依据页面摘要间的相关性而非比较搜索键。其结构由LLM为每个文档自动发现,而非手动设计。在NarrativeQA数据集上的初步评估中,该索引达到了与最强基准DocETL相当的准确率,而回答问题的成本降低了40%。

英文摘要

Long documents such as clinical records, legal contracts, and scientific papers are increasingly analyzed with large language models (LLMs). Naturally, feeding the full document to the model for every question can eventually become slow, expensive, prone to hallucination, and it reuses no work across questions. We explore an indexing-based solution for document analysis and propose a hierarchical plain-text index that is built once per document and consulted by subsequent queries. Inspired by the classic B+ tree, the index organizes a document into pages arranged from general summaries at the root to specific ones at the leaves, but it departs from the B+ tree in three ways suited to LLM access: content lives at every level, the index is plain text rather than attribute values, and navigation follows relevance between page summaries rather than comparing a search key. Its structure is discovered per document by the LLM rather than being hand-designed. In a preliminary evaluation on the NarrativeQA dataset, the index reaches accuracy comparable to DocETL, the strongest baseline, while being 40\% cheaper to answer questions.

Comments4 pages, 3 figures. VLDB 2026 PhD Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑