先结构后查询:支持对非结构化文档进行精确分析查询
Structure then Query: Enabling Precise Analytical Queries over Unstructured Documents
浏览论文内容
中文总结 AI 辅助
AnnoIndex通过注释索引与结构化查询引擎,解决非结构化文档精确分析查询难题,在三个数据集上平均F1达0.87,优于现有基线方法。
中文摘要 AI 辅助
非结构化文档构成了企业和网络数据的主体。随着大语言模型(LLMs)的快速发展,研究人员已开始构建可像操作数据库一样分析非结构化文本文档的数据系统。然而,由于主流检索方法仍依赖基于向量相似度的模糊匹配,准确获取信息并执行结构化分析与推理仍是一项重大挑战。为解决这些局限,AnnoIndex引入了两个核心基础组件。第一个是注释索引(Annotation Index):系统使用名为SchemaLoop的模块从原始语料库自动创建分层注释模式,再用轻量级语言模型提取特定值,将分散的非结构化文本转化为物化的结构化索引,支持低成本过滤与查询;该注释索引避免了向量相似度的黑盒匹配,将属性提取成本从在线查询分摊到一次性构建阶段。第二个创新是结构化查询引擎(Structured Query Engine):它基于SQL扩展将用户问题编译为执行计划,先利用注释索引进行精确文档过滤,再按成本升序逐步应用提取操作,仅对语料库中需要深度语义理解的最小剩余部分调用LLMs;提取的属性会合并到注释索引中,降低后续查询的成本。在三个真实世界数据集上的实验表明,AnnoIndex始终优于当前最先进的基线方法,取得了最高的平均F1分数(0.87),同时在复杂多跳连接与渐进式推理查询上保持稳健性能。
英文摘要
Unstructured documents constitute the majority of enterprise and web data. With the rapid development of large language models(LLMs), researchers have started to build data systems that analyze unstructured textual documents like operating on databases. However, because mainstream retrieval methods still relies on fuzzy matching based on vector similarity, accurately obtaining information and performing structured analysis and reasoning remains a major challenge. To address these limitations, AnnoIndex introduces two core fundamental components. The first is Annotation Index. The system uses a module called SchemaLoop to automatically create hierarchical annotation schemas from the raw corpus, and then uses lightweight language model to extract specific values. It turns scattered unstructured text into a materialized, structured index that enables low-cost filtering and querying. The annotation index avoids the black-box matching of vector similarity and amortizes attribute extraction costs from online queries to a one-time build. The second innovation is a Structured Query Engine. It compiles user questions into execution plans based on SQL extension. It first uses the Annotation Index for precise documents filtering, then gradually applies extraction operations in ascending order of cost, resorting to LLMs only for the remaining minimal fraction of the corpus that require deep semantic understanding. The extracted attributions are merged into the annotation index, reducing the cost of future queries. Experiments on three real-world datasets demonstrate that AnnoIndex consistently outperforms state-of-the-art baselines, achieving the highest average F1 score (0.87) while maintaining robust performance on complex multi-hop join and progressive reasoning queries.