arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19491cs.DBcs.AI

高效链接非结构化数据以支持多步推理

Efficiently Linking Unstructured Data for Multi-step Reasoning

Jiaming Liang, Haydn Jones, Jacob R. Gardner, Mark Yatskar, Zachary Ives

首次发表
浏览论文内容

中文总结 AI 辅助

针对多步推理中的非结构化数据链接问题,提出DASE查询引擎,通过多步推理模型、稀疏索引和协同执行层,实现比现有基线快6-46倍的检索,并提升下游LLM评估质量与成本效益。

中文摘要 AI 辅助

现代大语言模型(LLM)和AI智能体日益支持整合非结构化来源证据的数据工程工作流。此类流水线通常先进行数据检索、整合和排序,再执行更复杂的智能体推理或操作,例如用于科学发现。这些工作流中的核心检索问题需联合执行多属性过滤、多向量搜索、精确关系连接以及阈值化嵌入相似度连接。针对给定的规划查询和单调评分函数,我们的DASE查询引擎构建并排序候选证据元组。它包含:(i)一个基于结构化谓词、多向量和关系链接的多步推理查询模型;(ii)SemJI,一种用于稀有近邻对的稀疏物化嵌入相似度连接索引;(iii)一个协同设计的执行层,结合了谓词感知的近似最近邻(ANN)遍历、批量访问和基于阈值的分数聚合。在科学发现工作负载上,DASE在可比较的召回率下,比强大的关系数据库管理系统(RDBMS)、重排序和向量数据库基线快6倍到46倍地检索多步推理查询的候选证据;对于需要语义算子后处理的任务,DASE作为高召回率预过滤器,使下游LLM评估既更便宜又更准确——例如,在SemBench E-Commerce上,它将BigQuery的质量从0.67提升到0.80,同时将成本从2.42美元降至0.54美元。

英文摘要

Modern LLMs and AI agents increasingly support data engineering workflows that integrate evidence from unstructured sources. Such pipelines typically do data retrieval, integration, and ranking before proceeding to more complex agentic reasoning or actions, e.g., for scientific discovery. The core retrieval problem in these workflows jointly executes multi-attribute filtering, multi-vector search, exact relational joins, and thresholded embedding-similarity joins. Given a planned query and monotone scoring function, our DASE query engine constructs and ranks candidate evidence tuples. It comprises (i) a multi-step reasoning query model over structured predicates, multiple vectors, and relational links; (ii) SemJI, a sparse materialized embedding-similarity join index for rare near-neighbor pairs; and (iii) a co-designed execution layer that combines predicate-aware ANN traversal, batched access, and threshold-based score aggregation. On scientific-discovery workloads, DASE retrieves candidate evidence for multi-step reasoning queries 6x to 46x faster than strong RDBMS, rerank, and vector-database baselines at comparable recall; and for tasks that require semantic-operator post-processing, DASE acts as a high-recall prefilter that makes downstream LLM evaluation both cheaper and more accurate -- e.g., on SemBench E-Commerce it improves BigQuery quality from 0.67 to 0.80 while cutting cost from $2.42 to $0.54.

发表机构

  • University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑