发表机构
Zleap AI(智量科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出SAG(SQL检索增强生成)架构,通过查询时动态超边实现结构化检索,在HotpotQA等多跳问答基准上性能优于基线,为LLM智能体处理组织知识提供了新方案。
AI 中文摘要
尽管检索增强生成(RAG)已被证明能有效让大语言模型(LLM)获取外部知识,但主流的密集检索实现仍在处理结构化约束和多跳推理方面存在固有局限。基于图的方法通过离线构建知识图谱来解决这一问题,但它们常导致语义碎片化、维护成本高且增量更新复杂。我们提出SAG(SQL检索增强生成),这是一种结构化检索架构,无需构建全局知识图谱,而是将文档组织成事件-实体索引。SAG将每个文本块表示为语义完整的事件及其关联实体,形成潜在超边,保留n元关系而不将其分解为三元组。在查询时,SAG将共享实体视为连接键来关联相关文本块,动态生成查询范围内的事件邻域,且所有证据始终保持原始文本块。在HotpotQA、2WikiMultiHopQA和MuSiQue上的实验表明,SAG在所有基准测试中均取得了最佳的检索和端到端问答性能,且收益随推理链复杂度增加而扩大。在对多跳证据链要求最高的MuSiQue上,SAG的Recall@5达到80.36%,比最强基线高出11.52个百分点。本研究为能让LLM智能体检索和推理不断增长的组织知识的知识基础设施奠定了基础。
英文摘要
While retrieval-augmented generation (RAG) has proven effective at giving LLMs access to external knowledge, mainstream dense-retrieval implementations remain inherently limited in handling structured constraints and multi-hop reasoning. Graph-based methods address this by constructing knowledge graphs offline, but they often fragment semantics, incur high maintenance, and complicate incremental updates. We propose SAG (SQL-Retrieval Augmented Generation), a structured retrieval architecture that organizes documents into an event-entity index without building a global knowledge graph. SAG represents each chunk as a semantically complete event paired with its entities, forming a latent hyperedge that preserves n-ary relations without decomposing them into triples. At query time, SAG treats shared entities as join keys to connect related chunks. This dynamically yields a query-scoped neighborhood of events, and yet every piece of evidence remains the original chunk throughout. Experiments on HotpotQA, 2WikiMultiHopQA, and MuSiQue show that SAG achieves the best retrieval and end-to-end QA performance on every benchmark, with gains that widen as reasoning-chain complexity increases. On MuSiQue, where multi-hop evidence chaining is most demanding, SAG reaches 80.36% Recall@5, outperforming the strongest baseline by 11.52 points. This work paves the way for knowledge infrastructure that enables LLM agents to retrieve and reason over continually growing organizational knowledge.
CommentsThis work was intended as a replacement of arXiv:2606.15971 and any subsequent updates will appear there