arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

G-ReAct:基于结构-状态协同演化的图引导深度搜索

G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution

Shaoxiong Yang, Mengyuan Zhang, Shaojun Lin, Chao Li, Wei Liu, Kun Shao, Jian Luan

arXiv 2608.01324首次发表:更新:

AI 中文总结

提出图引导深度搜索框架G-ReAct,通过结构-状态协同演化解决现有LLM深度搜索的上下文遗忘等问题,仅用少量轨迹微调即提升模型在BrowseComp-ZH、XBench上的准确率,且推理时可增强现有LLM性能。

AI 中文摘要

深度搜索已成为大型语言模型(LLMs)解决开放域复杂任务的基础能力。然而,现有方法通常依赖线性序列推理进行轨迹生成和推理,难以在长 horizon 多跳搜索中一致保留中间状态和约束,因此常出现上下文遗忘、搜索漂移和探索效率低下的问题。为解决这些局限,我们提出 G-ReAct,一种深度搜索的推理框架,将推理组织为「固定拓扑查询图上的状态演化」。演化的图状态明确跟踪搜索进度并指导后续决策,将文本历史驱动的探索性搜索转变为显式约束下的图引导推理。G-ReAct 支持训练和推理:它生成高质量深度搜索轨迹用于监督微调,并在推理时为搜索提供结构化指导,无需额外微调。实验表明,仅用 1900 条生成轨迹微调,Qwen3-30B-A3B-Thinking-2507 在 BrowseComp-ZH 上准确率达 52.6%,在 XBench 上达 79.0%,优于在大得多的数据集上训练的可比开源方法,包括强化学习增强的方法。此外,在推理时应用 G-ReAct,可持续提升现有强大 LLMs 在深度搜索任务上的性能。我们将公开发布所有代码和模型权重。

英文摘要

Deep search has become a fundamental capability of large language models (LLMs) for solving open-domain complex tasks. However, existing approaches typically rely on linear sequential reasoning for both trajectory generation and inference, making it difficult to consistently preserve intermediate states and constraints throughout long-horizon multi-hop search. Consequently, they often suffer from context forgetting, search drift, and inefficient exploration. To address these limitations, we propose $\textbf{G-ReAct}$, a reasoning framework for deep search that organizes reasoning as $\textbf{state evolution over a fixed-topology query graph}$. The evolving graph state explicitly tracks search progress and guides subsequent decisions, transforming exploratory search driven by textual history into graph-guided reasoning under explicit constraints. G-ReAct supports both training and inference: it generates high-quality deep-search trajectories for supervised fine-tuning and provides structured guidance for inference-time search without additional fine-tuning. Experiments demonstrate that with only 1.9K generated trajectories for fine-tuning, Qwen3-30B-A3B-Thinking-2507 achieves $52.6\%$ accuracy on BrowseComp-ZH and $79.0\%$ on XBench, outperforming comparable open-source methods trained on substantially larger datasets, including RL-enhanced methods. Furthermore, when applied at inference time, G-ReAct consistently improves the performance of existing strong LLMs on deep-search tasks. We will publicly release all code and model weights.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑