arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26604cs.CL

WikiLoop:基于下游反馈联合学习构建与智能体原生Wiki导航

WikiLoop: Jointly Learning to Build and Navigate Agent-Native Wikis with Downstream Feedback

Haoliang Ming, Feifei Li, Wenhui Que

首次发表
浏览论文内容

中文总结 AI 辅助

WikiLoop是反馈耦合框架,以Qwen3.5-9B为主干,联合学习构建与导航智能体原生Wiki,在AuthTrace等数据集上提升答案正确性,整合两种角色能力且无需特定数据集训练。

中文摘要 AI 辅助

知识库构建与查询通常被孤立优化:检索增强智能体在固定的外部维护索引上运行,而构建过程无法从下游使用中获得反馈。我们提出WikiLoop,这是一种反馈耦合框架,用于联合学习构建和导航智能体原生Wiki——一种专为机器导航设计的持久链接页面知识库。角色条件共享策略支持两种接口:导航器从Wiki中检索证据以回答查询,构建器提出结构化编辑,通过下游导航进行评估。导航器遵循“先充分后高效”的目标,仅在收集完整证据后才应用检索成本惩罚。构建器从效用差异中学习:冻结的导航器通过候选编辑对下游性能的变化进行评分,同时设置保护惩罚以阻止对不相关查询的性能下降。训练过程结合了特定角色的顺序优化,以及最终针对角色同质批次的联合阶段。以Qwen3.5-9B为通用主干,WikiLoop在AuthTrace上达到62.6的综合答案正确性,比LLM-Wiki base高出6.3个百分点,在多文档查询上的增益最大。控制对比验证了两个目标的预期效果,且学习到的编辑对保留的导航器仍然有用。配对对比表明,最终的共享策略在很大程度上保留了两种特定角色的能力,与相应的专家基线相比,导航器和端到端答案正确性分别提高了0.4个百分点,并将两种接口整合到一个模型中。在未进行特定数据集训练的情况下,WikiLoop在HotpotQA和MuSiQue上也优于同主干的LLM-Wiki base。

英文摘要

Knowledge-base construction and querying are typically optimized in isolation: retrieval-augmented agents operate over a fixed, externally maintained index, whereas construction receives no signal from downstream use. We present WikiLoop, a feedback-coupled framework that jointly learns to build and navigate an agent-native Wiki, a persistent linked-page knowledge base designed for machine navigation. A role-conditioned shared policy supports two interfaces: a Navigator retrieves evidence from the Wiki to answer queries, and a Builder proposes structured edits evaluated through downstream navigation. The Navigator follows a sufficiency-before-efficiency objective that applies retrieval-cost penalties only after full evidence has been collected. The Builder learns from utility differences: a frozen Navigator scores each candidate edit by its change in downstream performance, while a guard penalty discourages regressions on unrelated queries. Training combines sequential role-specific optimization with a final joint stage over role-homogeneous batches. With Qwen3.5-9B as the common backbone, WikiLoop reaches 62.6 aggregate Answer Correctness on AuthTrace, 6.3 points above LLM-Wiki, base, with the largest gains on multi-document queries. Controlled comparisons support the intended effects of both objectives, and the learned edits remain useful to a held-out Navigator. Paired comparisons indicate that the final shared policy largely retains both role-specific capabilities, improves Navigator and end-to-end Answer Correctness by 0.4 points relative to the corresponding specialist references, and consolidates both interfaces into one model. Without dataset-specific training, WikiLoop also improves over the same-backbone LLM-Wiki, base on HotpotQA and MuSiQue.

补充信息

↑