SearchWiki:学习构建与导航知识Wiki以支持主动信息检索
SearchWiki: Learning to Build and Navigate Knowledge Wikis for Active Information Seeking
浏览论文内容
中文总结 AI 辅助
该研究提出SearchWiki框架,将语料库合成分层可导航Wiki,训练WikiResearcher-9B智能体经强化学习优化导航策略,在多基准测试中表现优于同规模模型,证实结构化语料库学习式导航优于扁平检索。
中文摘要 AI 辅助
基于扁平检索增强生成的方法将语料库视为块的集合,忽略了文档层级与跨文档结构。我们提出SearchWiki,一个可将语料库合成为分层、带类型、可导航的Wiki的框架,并训练智能体WikiResearcher-9B通过多轮工具使用进行信息检索。该Wiki将知识组织为三层:文档概述、跨文档主题页、页面级源记录,使初始检索未命中时可逐步优化检索。我们采用带多分量奖励函数的在线策略强化学习优化智能体的导航策略,奖励函数平衡答案正确性、检索质量与轨迹效率。在ViDoRe-V3(8个领域)、FinanceBench及记忆基准(LoCoMo、LongMemEval、PersonaMem-v2)上的评估显示,作为经强化学习调优的Qwen 9B模型,WikiResearcher-9B显著优于同规模未训练基线,且超过或匹配更大的外部模型。SearchWiki与WikiResearcher-9B配对表明,对结构化语料库的学习式导航是扁平检索的更优替代方案。
英文摘要
Flat retrieval-augmented generation treats a corpus as a bag of chunks, discarding document hierarchy and cross document structure. We introduce SearchWiki, a harness framework that synthesizes a corpus into a hierarchical, typed, navigable wiki and trains an agent, WikiResearcher-9B, to retrieve information through multi-turn tool use. The wiki organizes knowledge into three layers - document overviews, cross- document topic pages, and page-level source records; enabling progressive refinement of retrieval when initial lookup misses. We optimize the agent's navigation policy with on-policy reinforcement learning with a multi-component reward function balancing answer correctness, retrieval quality and trajectory efficiency. Evaluation on ViDoRe-V3 (8 domains), FinanceBench, and memory benchmarks (LoCoMo, LongMemEval, PersonaMem-v2) shows that WikiResearcher- 9B which is our RL-tuned Qwen 9B model, significantly outperforms same-size untrained baselines and exceeds or matches larger external models. SearchWiki paired with WikiResearcher-9B demonstrates that learned navigation over structured corpora is a superior alternative to flat retrieval.
发表机构
- IBM(国际商业机器公司)
机构由 AI 辅助整理,请以论文原文为准。