CoEvoKG:用自演进搜索智能体协同演进知识图谱
CoEvoKG: Co-Evolving Knowledge Graphs with Self-Evolving Search Agents
浏览论文内容
中文总结 AI 辅助
CoEvoKG框架将知识图谱作为训练任务源与智能体演进的持久证据记忆,联合训练任务生成器与搜索智能体,形成自演进与知识积累闭环,在六个问答基准上提升了三款骨干模型的准确率。
中文摘要 AI 辅助
大型语言模型可通过强化学习提升搜索智能体性能,但现有自博弈智能体在反复生成任务时,会丢弃成功搜索过程中获得的知识。本文提出CoEvoKG框架,将知识图谱转化为可验证训练任务的来源,同时作为智能体演进的持久证据记忆。CoEvoKG联合训练任务生成器与搜索智能体:生成器从知识图谱采样的实体链中创建多跳问题,智能体根据答案正确性及实体路径有图谱证据支持的搜索轨迹获得奖励。搜索成功时,CoEvoKG会验证并去重检索到的证据,再将其写回对应图谱节点与边,后续轮次复用该丰富后的图谱进行任务生成与奖励计算,形成模型自演进与知识积累的闭环。在六个问答基准(NQ、TriviaQA、PopQA、HotpotQA、2WikiMultiHopQA、Bamboogle)及三个骨干模型上的实验显示,CoEvoKG使Qwen2.5-3B-Instruct、Qwen2.5-7B-Instruct、Llama-3.1-8B-Instruct的宏观平均准确率较对应基础模型分别提升11.2、10.1、11.6个百分点;在匹配的训练预算下,CoEvoKG较三个骨干模型的竞争性自博弈基线与搜索智能体RL基线,宏观平均准确率还提升2.6至3.7个百分点。代码可在该https网址获取。
英文摘要
Large language models can improve with reinforcement learning for search agents, yet existing self play agents repeatedly generate tasks while discarding the knowledge gained during successful searches. We introduce CoEvoKG, a framework that turns a knowledge graph into both a source of verifiable training tasks and a persistent evidence memory for agent evolution. CoEvoKG jointly trains a task generator and a search agent: the generator creates multihop questions from entity chains sampled from the knowledge graph, while the agent learns from rewards for answer correctness and search trajectories whose entity paths are supported by graph evidence. When a search succeeds, CoEvoKG verifies and deduplicates the retrieved evidence, then writes it back to the corresponding graph nodes and edges. Future rounds reuse this enriched graph for task generation and reward computation, closing the loop between model self evolution and knowledge accumulation. Experiments on six QA benchmarks (NQ, TriviaQA, PopQA, HotpotQA, 2WikiMultiHopQA, and Bamboogle) with three backbone models show that CoEvoKG improves macro average accuracy over the corresponding base models by +11.2, +10.1, and +11.6 points on Qwen2.5-3B-Instruct, Qwen2.5-7B-Instruct, and Llama-3.1-8B-Instruct, respectively. Under matched training budgets, CoEvoKG further improves over competitive self play baselines and RL baselines for search agents by +2.6 to +3.7 macro average points across the three backbones. Code is available at https://github.com/lazzy1225/CoEvoKG.