发表机构
Beijing Academy of Artificial Intelligence (BAAI)(北京通用人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对深度研究中发现与验证答案的不对称性,提出递归自我改进智能体AREX。它通过内部研究与外部自我改进循环交替工作,学习自主上下文更新工具,经训练后在多个基准测试中显著优于同等规模基线,与参数更多模型竞争。
AI 中文摘要
深度研究要求智能体找到能共同满足多个约束的答案。发现此类答案成本高昂,而验证候选答案通常可分解为易于处理的按约束检查。这种发现与验证的不对称表明,研究智能体不应只是简单地搜索更长时间,而应通过验证中间结果并利用部分验证状态来指导后续改进,从而递归地改进当前答案。我们引入了AREX,这是一族递归自我改进(RSI)深度研究智能体。AREX在收集证据并构建临时答案的内部研究循环和按约束审核答案、识别未解决的主张并开展有针对性的后续研究的外部自我改进循环之间交替。为了长期维持RSI,AREX学习一种自主上下文更新工具,该工具将不断增长的交互历史压缩为一个紧凑的改进状态,保留已验证的证据和未解决的约束,而不依赖外部模型。我们通过智能体中期训练和长期强化学习在经过验证的数据合成任务和高质量轨迹上训练AREX。为了减轻长期学习过程中稀疏的最终奖励问题,我们强调获取决定性证据或纠正错误研究方向的关键步骤。我们实例化了一个密集的4B模型和一个122B - A10B专家混合模型。在浏览比较、广泛搜索深度搜索问答、人类最后考试(HLE)等推理和工具使用基准测试中,AREX显著优于同等规模的基线,并且与使用更多激活参数的模型相比仍具有竞争力。
英文摘要
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermediate results and using the partially verified state to guide subsequent refinement. We introduce AREX, a family of Recursively Self-Improving (RSI) deep research agents. AREX alternates between an inner research loop that gathers evidence and constructs a provisional answer, and an outer self-improvement loop that audits the answer constraint-wise, identifies unresolved claims, and launches targeted follow-up research. To sustain RSI over long horizons, AREX learns an autonomous context-update tool that compresses growing interaction history into a compact improvement state preserving verified evidence and unresolved constraints, without relying on an external model. We train AREX on verified synthetic tasks and high-quality trajectories through agentic mid-training and long-horizon reinforcement learning. To mitigate sparse final rewards during long horizon learning, we emphasize key steps where decisive evidence is acquired or erroneous research directions are corrected. We instantiate a dense 4B model and a 122B-A10B Mixture-of-Experts model. Across BrowseComp, WideSearch, DeepSearchQA, Humanity's Last Exam (HLE), and other reasoning and tool-use benchmarks, AREX substantially outperforms comparable-scale baselines and remains competitive with models using substantially more activated parameters.