STITCH-RAG:面向多跳检索增强生成的主题超图时空影响追踪
STITCH-RAG: Spatio-Temporal Influence Tracing over Topic Hypergraphs for Multi-Hop Retrieval-Augmented Generation
- School of Statistics And Data Science, Guangdong University of Finance & Economics(广东财经大学统计与数据科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出STITCH-RAG超图框架,通过半合并主题超图、时空影响桥接传播和连续PPR先验,在多跳检索中提升证据连接与生成质量,在HotpotQA和2WikiMultiHopQA上取得最优结果。
AI中文摘要:
多跳检索增强生成需要检索器连接分布在多个文档中的证据,同时保持简洁、忠实的生成上下文。现有索引存在两个互补的缺陷:基于分块的RAG可能破坏跨段落证据链,而无标签的成对投影(不包含生成主题来源)无法同时保留主题级共现关系和每次出现的实体描述。我们提出STITCH-RAG,一个包含三个耦合组件的超图框架。首先,半合并主题超图将多实体共现编码为主题摘要超边,同时保留由规范名称等价性链接的每分块实体状态。其次,时空影响桥接传播(STIBP)结合主题空间传播与确定性分块索引链接,在频率自适应衰减下跨越名称等价状态。第三,连续的STIBP分数取代局部个性化PageRank(PPR)中的二元实体匹配种子。我们刻画了该先验比二元先验将更多PPR质量分配给真实证据的条件。在报告协议下,STITCH-RAG在HotpotQA和2WikiMultiHopQA上取得了所比较方法中最高的Contain-Acc和LLM-Acc点估计,并在标准化检索比较中取得比所包含方法更高的Recall@8。混合领域基准上的结果仍作为辅助的基于偏好的证据,因为只有LLM评判的准确率可用。
英文摘要:
Multi-hop retrieval-augmented generation requires a retriever to connect evidence distributed across documents while preserving a concise, faithful generation context. Existing indexes leave two complementary gaps: chunk-based RAG can break cross-passage evidence chains, whereas an unlabeled pairwise projection without generating-topic provenance cannot jointly preserve topic-level co-participation and per-occurrence entity descriptions. We propose STITCH-RAG, a hypergraph-based framework with three coupled components. First, a semi-merged topic hypergraph encodes multi-entity co-participation as topic-summary hyperedges while retaining per-chunk entity states linked by canonical-name equivalence. Second, spatio-temporal influence bridging propagation (STIBP) combines topic-space propagation with deterministic chunk-index linkage across name-equivalent states under frequency-adaptive decay. Third, continuous STIBP scores replace binary entity-match seeds in localized Personalized PageRank (PPR). We characterize the condition under which this prior assigns more PPR mass to ground-truth evidence than a binary prior. Under the reported protocol, STITCH-RAG attains the highest reported Contain-Acc and LLM-Acc point estimates among the compared methods on HotpotQA and 2WikiMultiHopQA, and higher Recall@8 than the methods included in the standardized retrieval comparison. Results on the mixed-domain benchmark remain auxiliary preference-based evidence because only LLM-judged accuracy is available.