arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

质疑问题:推理模型中自我进化的持续维持

Questioning the Questions: Sustaining Self-Evolution in Reasoning Models

Jinyuan Li, Chengsong Huang, Langlin Huang, Donghong Cai, Shiping Gao, Yuyi Yang, Jiaxin Huang

arXiv 2610.04299首次发表:更新:

发表机构

Washington University in St. Louis; University of Michigan, Ann Arbor(圣路易斯华盛顿大学; 密歇根大学安娜堡分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对推理模型自我进化中性能崩溃问题,提出R-Quest方法,通过问题有效性与新颖性反馈维持进化,在12个基准上取得最优平均性能,十轮进化中稳定提升并超越R-Zero 17.32点。

AI 中文摘要

自我进化推理模型从自身生成的问题中学习,然而重复的自我训练可能导致性能崩溃。在本文中,我们研究了为什么性能在连续多轮中会恶化,以及如何维持自我进化。我们的分析识别出自我生成问题中两个反复出现的质量问题:无效问题和同一数学问题的重复变体。首先,无效问题在后续轮次中变得更加普遍,而答案一致性过滤进一步增加了其在训练数据中的比例。其次,基于词汇相似性的现有问题多样性控制可能遗漏以不同方式表达的数学等价问题,这导致在后续训练轮次中出现问题多样性崩溃。基于这些发现,我们引入了R-Quest,它利用问题有效性和新颖性反馈来引导自我进化。我们首先训练求解器识别并拒绝无效问题,然后利用其判断来引导提问者奖励并过滤求解器训练数据。为了避免问题重复,我们使用冻结的基础模型来比较采样的问题对并提供新颖性反馈。实验上,我们的方法在两个模型家族中的数学推理、通用领域推理和代码生成等12个基准上始终取得最高的平均性能。此外,R-Quest在十轮自我进化中保持稳定的性能提升,在最后一轮达到峰值,并比R-Zero高出17.32个百分点。

英文摘要

Self-evolving reasoning models learn from their own generated questions, yet repeated self-training can lead to performance collapse. In this paper, we investigate why performance deteriorates over successive rounds and how to sustain self-evolution. Our analysis identifies two recurring quality problems in self-generated questions: invalid questions and repeated variants of the same mathematical questions. First, invalid questions become more prevalent across rounds, and answer-consistency filtering further increases their proportion in training data. Second, existing question diversity controls based on lexical similarity can miss mathematically equivalent questions expressed in different ways, which leads to question diversity collapse in later training rounds. Building on these findings, we introduce R-Quest, which uses question validity and novelty feedback to guide self-evolution. We first train the solver to recognize and reject invalid questions, then use its judgments to guide questioner rewards and filter solver training data. To avoid question repetition, we use a frozen base model to compare sampled question pairs and provide novelty feedback. Empirically, our method consistently achieves the highest average performance on 12 benchmarks in mathematical reasoning, general-domain reasoning, and code generation across two model families. Additionally, R-Quest maintains stable performance gains over ten rounds of self-evolution, peaking in the final round and outperforming R-Zero by 17.32 points.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑