Search-GRT:面向复杂问答任务优化的搜索智能体引导检索训练
Search-GRT: Guided Retrieval Training of Search Agents to Optimize for Complex Question Answering
浏览论文内容
中文总结 AI 辅助
针对LLM搜索的奖励稀疏问题,提出GRT方法,通过真实信息限制检索过程,在多跳问答等任务上实现超40%性能提升并提高训练效率。
中文摘要 AI 辅助
大型语言模型(LLMs)有效利用搜索引擎仍是一项重大挑战,尤其在复杂多跳问答(MHQA)任务中。这些任务要求模型将问题分解为子查询、检索相关信息并从多源合成答案,早期检索不佳常导致连锁错误。强化学习(RL)虽有望提升LLMs的搜索能力,但训练中常面临奖励稀疏问题,阻碍模型有效学习。为应对这些挑战,我们提出引导检索训练(Guided Retrieval Training,GRT)这一新方法,通过在RL训练中利用真实信息限制检索过程来提升搜索智能体性能。通过聚焦精心筛选的相关文档集,GRT为模型提供更强的学习信号,缓解奖励稀疏问题,提升其生成准确子查询和合成正确答案的能力。实验结果表明,在各类问答(QA)任务中,GRT均实现了对现有方法(如Search-R1)的持续性能提升;尤其在MHQA任务中,GRT性能提升超40%,且能以更少训练步骤获得更优QA性能,提升了训练效率。
英文摘要
The effective use of search engines by large language models (LLMs) remains a significant challenge, particularly in complex, multi-hop question-answering (MHQA) tasks. These tasks require the model to decompose questions into subqueries, retrieve relevant information, and synthesize answers from multiple sources, often leading to cascading errors due to poor retrieval in early stages. Reinforcement learning (RL) has shown promise in improving LLMs' search capabilities, but it often suffers from sparse rewards during training, hindering the model's ability to learn effectively. To address these challenges, we introduce Guided Retrieval Training (GRT), a novel method that improves the performance of a search agent by restricting the retrieval process during RL training using ground truth information. By focusing on a curated set of relevant documents, GRT provides the model with a stronger learning signal, mitigating the problem of sparse rewards and improving its ability to generate accurate subqueries and synthesize correct answers. Our experimental results demonstrate that GRT achieves consistent performance improvements over existing methods, such as Search-R1, across a wide range of question-answering (QA) tasks. Notably, GRT excels in MHQA tasks, achieving over 40% improvements in performance. Additionally, GRT enhances training efficiency by achieving better QA performance with fewer training steps.
发表机构
- AI Center-Mountain View, Samsung Electronics(三星电子山景城AI中心)
机构由 AI 辅助整理,请以论文原文为准。