发表机构
KT Corporation(KT公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Mosaic提出无需训练的逐查询策略自适应框架,通过LLM分析器动态调整图探索策略,在GraphRAG-Bench上显著提升答案正确率并减少计算开销。
AI 中文摘要
图检索增强生成(GraphRAG)能够连接分布在语料库图中的证据,但大多数系统在不同查询间使用大体共享的探索流程。这造成了结构性错配:直接事实可能需要紧凑的局部邻域,比较类问题需要多个目标的均衡覆盖,而间接问题可能需要通过弱关联连接器的更深路径。我们提出Mosaic,一个无需训练的框架,将GraphRAG检索表述为逐查询控制问题。一个LLM分析器将查询特定的证据需求转化为一个受限策略,涵盖种子选择、图遍历、停止和证据选择,而语料库图、索引、评分函数、接地过程和答案生成器保持共享。在GraphRAG-Bench上,Mosaic在Medical上达到查询加权答案正确率76.97,在Novel上达到64.33,比此前报告的最强总体结果分别提升5.13和4.43个百分点。在Medical上,它达到95.1的证据召回率和86.1的上下文相关性。在相同图和生成器上的受控比较表明,没有固定的窄、中或宽策略始终最优;Mosaic比最强的规范固定策略提升9.96个百分点。相对于固定宽策略,它评估的路径减少81.9%,保留的证据项减少47.2%。在HotpotQA、MuSiQue和2WikiMultiHopQA上的迁移实验进一步表明,该策略接口可在无需基准特定检索器训练的情况下应用。
英文摘要
Graph Retrieval-Augmented Generation (GraphRAG) can connect evidence distributed across a corpus graph, but most systems use largely shared exploration procedures across queries. This creates a structural mismatch: direct facts may need compact local neighborhoods, comparisons need balanced coverage of multiple targets, and mediated questions may require deeper paths through weakly related connectors. We present Mosaic, a training-free framework that formulates GraphRAG retrieval as a per-query control problem. An LLM analyzer converts query-specific evidence requirements into a bounded policy over seed selection, graph traversal, stopping, and evidence selection, while the corpus graph, indexes, scoring functions, grounding procedure, and answer generator remain shared. On GraphRAG-Bench, Mosaic achieves query-weighted Answer Correctness of 76.97 on Medical and 64.33 on Novel, improving over the strongest previously reported overall results by 5.13 and 4.43 points. On Medical, it reaches 95.1 Evidence Recall and 86.1 Context Relevancy. Controlled comparisons on an identical graph and generator show that no fixed narrow, medium, or wide policy is consistently optimal; Mosaic improves by 9.96 points over the strongest canonical fixed policy. Relative to Fixed Wide, it evaluates 81.9% fewer paths and retains 47.2% fewer evidence items. Transfer experiments on HotpotQA, MuSiQue, and 2WikiMultiHopQA further show that the policy interface can be applied without benchmark-specific retriever training.
Comments13 pages, 2 figures, 11 tables. Technical Report