发表机构
Rutgers University; University of California, San Diego; University of Michigan; McGill University; King Fahd University of Petroleum and Minerals(罗格斯大学; 加州大学圣迭戈分校; 密歇根大学; 麦吉尔大学; 法赫德国王石油矿产大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对自进化搜索智能体中的共作弊问题,提出CrossFit方法,通过交叉拟合反馈隔离源文档,显著降低虚假一致性并提升下游搜索性能。
AI 中文摘要
自进化搜索智能体通过联合优化一个生成问题的提议器和一个解答问题的求解器,来构建自己的训练课程。这种闭环引入了一种我们称之为“共作弊”的失败模式:提议器和求解器在共享错误上日益达成一致,导致内部奖励提升,但外部正确性却没有相应的提高。对源证据的事后审计显示,随着自进化的连续轮次推进,共作弊现象愈发严重,伪标签的正确性停滞甚至下降,即便循环内的训练信号在改善。最直接的缓解措施是在训练前验证提议:我们引入了多样本验证(MSV),该方法对同一模型分别在有源和无源条件下各查询三次,以决定任务是否准入并替换不可靠的伪标签。MSV部分减少了虚假一致性,但留下了大量的残余共作弊,并且每个候选需要额外的六次标签生成。这些局限性促使我们提出了主要方法CrossFit:它将提议器的源文档分为A组和B组;从A组生成的问题由仅在B组上训练的辅助求解器评分,反之亦然。交叉拟合的一致性决定了提议器的奖励,因此同源的伪标签无法通过反馈求解器重现,而原始求解器的更新规则保持不变。使用Qwen3.5-4B和Qwen3.5-9B重新运行循环,MSV将虚假一致性质量从6.1%降至5.7%,从8.8%降至7.2%,而CrossFit将其降至3.0%和3.7%。使用排除源的反馈重放相同的提议,进一步将虚假一致性降至0.4%和0.1%,从而将反馈的谱系与课程变化分离开来。在七个下游搜索基准测试中,CrossFit在4B和9B规模下,相较于标准耦合自进化,平均性能分别提升了8.8和8.4个百分点,相较于Search-R1分别提升了8.7和7.8个百分点。
英文摘要
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.
Comments21 pages. Equal contribution: Meijia Chen, Hao Li, Zheng Lu