Iris:攀登至搜索前沿
Iris: Climbing to the Search Frontier
浏览论文内容
中文总结 AI 辅助
本研究提出两款不同规模的开源搜索智能体Iris-mini和Iris-pro,采用SFT-RL攀登流程优化策略,在多个基准上取得同参数范围开源搜索智能体的最优结果,计划公开模型权重与完整方案。
中文摘要 AI 辅助
我们提出了Iris-mini和Iris-pro两款搜索智能体,分别在35B-A3B和397B-A17B规模下训练完成,同时公开了支撑它们的数据流水线与训练方案。任务从网页语料库的超链接结构反向构建而来:我们在从种子页面及其外链提炼出的实体图上生成多跳链,将所有非答案实体改写为描述性参考,使得任何线索都无法通过字符串匹配解析,仅保留参考模型以闭卷方式无法解决、但提供支撑证据后即可解答的问题。这些问题被转化为轨迹,在SFT前会经过轨迹级和步级两层过滤。随后,策略通过RL针对实时搜索进行优化,奖励评判器和观察总结器部署在训练集群内部,过长的rollout会在请求级别中断,并从已提交的前缀处恢复至下一步。我们将这两个阶段交替进行,该过程被称为SFT-RL攀登,将每一轮RL中最难解决且效率最高的rollouts返回给下一轮监督学习。由于推理时的上下文管理在这些基准上比系统间报告的大多数差异更有价值,我们对每个基准都在启用和未启用上下文管理的情况下进行评估,固定工具集、上下文限制和评判器。所有结果均来自单个ReAct智能体,无子智能体且无测试时验证。启用管理后,在BrowseComp、BrowseComp-ZH、DeepSearchQA和HLE上,两款模型分别达到82.2/84.8/86.9/52.3和88.6/85.1/92.9/56.4,是其参数范围内开源搜索智能体中整体表现最强的。我们计划发布模型权重以及完整的数据构建、训练和评估方案。
英文摘要
We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data pipeline and training recipe behind them. Tasks are reverse-constructed from the hyperlink structure of a web corpus: we author multi-hop chains over an entity graph distilled from a seed page and its out-links, rewrite every non-answer entity into a descriptive reference so that no clue can be resolved by string matching, and admit only questions that a reference model fails closed-book yet solves once the supporting evidence is supplied. These questions are then turned into trajectories, which are filtered at both the trajectory and the turn level before SFT. The policy is then optimized by RL against live search, with the reward judge and the observation summarizer served inside the training cluster, and with over-long rollouts interrupted at the request level and resumed from their committed prefix at the next step. We alternate the two stages in a procedure we call SFT-RL climbing, returning the hardest solved and most efficient rollouts of each RL round to the next supervised pass. Because inference-time context management is worth more on these benchmarks than most reported differences between systems, we evaluate every benchmark both with and without it, holding the tool set, the context limit, and the judge fixed. All results come from a single ReAct agent, with no sub-agents and no test-time verification. With management enabled, on BrowseComp, BrowseComp-ZH, DeepSearchQA, and HLE the two models reach $82.2/84.8/86.9/52.3$ and $88.6/85.1/92.9/56.4$, the strongest overall results among open-source search agents in their respective parameter ranges. We plan to release the model weights together with the complete recipe for data construction, training, and evaluation.
发表机构
- Shanghai AI Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。