发表机构
Shanghai Jiao Tong University; Shanghai Aircraft Manufacturing Co., Ltd.; Tongji University(上海交通大学; 上海飞机制造有限公司; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多跳检索增强生成中证据控制问题,提出DynaKRAG框架,将证据获取表述为状态条件控制,通过有效性层和控制器选择操作,实验表明该框架在多个基准测试中表现出色,证明协调检索等操作的益处。
AI 中文摘要
多跳检索增强生成(RAG)按顺序获取证据,每个新文档都可能揭示缺失的事实、桥梁实体、查询缺陷或足够的答案支持。现有方法提供了有用的操作,但通常在特定方法的管道或预定义的控制拓扑中组织它们。我们引入了DynaKRAG,将多跳证据获取表述为对原子证据操作的状态条件控制。在每个步骤中,有效性层构建可执行动作集,学习到的控制器选择下一个操作。实验结果表明,DynaKRAG在多个基准测试中优于最强的受控基线,证明了在不断演变的证据状态下协调检索、诊断和差距导向获取的好处。
英文摘要
Multi-hop retrieval-augmented generation (RAG) acquires evidence sequentially, with each document contributing supporting facts, bridge entities, query refinements, or sufficient evidence for answering. Evidence acquisition can involve iterative retrieval, query reformulation, evidence assessment, and sufficiency checking. We introduce DynaKRAG, a unified evidence-action framework that learns a shared state-conditioned policy for coordinating these operations. At each step, a deterministic validity layer constructs the executable action set, a learned continuation gate selects between answer generation and further evidence acquisition, and a learned advantage scorer ranks feasible evidence operations by their predicted gain relative to immediate answer generation. The selected operation updates the shared state and may enable additional operations. Across HotpotQA, 2Wiki, and MuSiQue with Qwen2.5-7B, GPT-4o-mini, and Llama-3.1-8B, DynaKRAG ranks first among the compared methods in both EM and F1 for all nine dataset--backbone pairs. Relative to matched-backbone baseline method, DynaKRAG improves F1 in every pair while achieving total-token efficiency gains of 10.1--34.3\% and retrieval-call efficiency gains of 15.1--43.4\%, establishing Pareto dominance under these measures. With Qwen2.5-7B, terminal evidence compression further improves answer quality across all three datasets while reducing the context passed to final answer generation by 54.4\%--71.5\%. These results demonstrate that unified, state-conditioned evidence control supports strong answer quality, efficient retrieval, and compact answer-generation contexts.