arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

少验证,多进化:为验证高效的机器学习进化智能体训练思想级评论模型

Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents

Jiamu Bai, Lizhu Zhang, Xin Yu, Yanhong Wu, Zellux Wang, Serena Li, Weiwei Li, Zhuokai Zhao, Lingzhou Xue, Kiwan Maeng, Xiangjun Fan, Bo Peng

arXiv 2610.08993首次发表:更新:

发表机构

Penn State University; Meta(宾夕法尼亚州立大学; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对机器学习进化智能体验证效率低的问题,提出训练思想级评论模型预测修改优劣,以筛选思想并优化验证资源分配,实验表明其优于前沿模型并提升解决方案质量与策略学习。

AI 中文摘要

随着大语言模型能力不断增强,自进化智能体能够应对包括机器学习人工智能(AI4ML)在内的挑战性任务。在AI4ML中,虽然经验验证可用,但通常需要计算成本高昂的模型训练和评估,限制了智能体进化的速度和规模。然而,验证效率仍未得到充分探索,且前沿模型在直接用作思想选择器时仅能提供有限的增益。我们通过专门的思想级评论模型来解决这一差距,该模型预测所提出的机器学习修改是否会优于当前解决方案,从而使智能体能够筛选思想并将验证资源集中在最有前景的候选方案上。我们通过Gemini-3.1-Pro合成的高质量评论进行监督微调来训练评论模型,随后使用GRPO进一步提高其预测准确性。实验上,我们的评论模型在静态思想评估中优于Gemini-3.1-Pro,且这些增益扩展到智能体推理、持续学习和策略训练。在推理时进化中,它们在相同验证预算下通过选择更有前景的思想来提高最终解决方案质量,并通过持续学习获得进一步增益。在策略训练中,它们充当学习到的奖励模型,将经验验证保留给不确定情况,并在相同验证资源下实现显著更多的策略更新。综合这些结果表明,思想级评论模型帮助机器学习智能体在有限验证预算下发现更好的解决方案并学习更强的提案策略。

英文摘要

As large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML). In AI4ML, while empirical verification is available, it often requires computationally costly model training and evaluation, limiting the speed and scale of agent evolution. Yet verification efficiency remains under-explored, and frontier models provide only limited gains when used directly as idea selectors. We address this gap with specialized idea-level critic models that predict whether a proposed ML modification will improve upon the current solution, allowing agents to screen ideas and concentrate verification resources on the most promising candidates. We train the critic models through supervised fine-tuning on high-quality critiques synthesized by Gemini-3.1-Pro, followed by GRPO to further improve their predictive accuracy. Empirically, our critic models outperform Gemini-3.1-Pro in static idea evaluation, and these gains extend to agent inference, continual learning, and policy training. During inference-time evolution, they improve final solution quality under the same verification budget by selecting more promising ideas, with further gains from continual learning. During policy training, they serve as learned reward models, reserving empirical verification for uncertain cases and enabling substantially more policy updates with the same verification resources. Together, these results show that idea-level critic models help ML agents discover better solutions and learn stronger proposal policies under limited verification budgets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑