arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14408cs.AI

通过成对验证器实现无奖励进化智能体

Reward-Free Evolving Agents via Pairwise Validator

Minghao Liu, Yu Wang, Jiayun Wang, Wei Wei

首次发表
浏览论文内容

中文总结 AI 辅助

研究提出用成对验证器取代标量奖励设计,集成到三个自我进化引擎中,有自适应焦点和软Elo两种变体。该方法在多智能体和两种工件基质上多数设置下达到或超全奖励基线,无需标记成本,是替代每步奖励设计的有效方案。

中文摘要 AI 辅助

一个自我进化的智能体循环会反复提出智能体(其提示模板或程序)的调整版本,并根据每次迭代的质量信号接受或拒绝该更改。设计该信号通常是项目中成本高昂的部分:可靠的标量奖励需要领域专业知识和标记示例,而这些示例的组装成本与智能体的基础任务一样高。我们建议用成对验证器取代接受/拒绝门处的标量:一个冻结的语言模型,给定父候选和子候选,返回关于哪个更好的二元判断。由于其对比性质,成对判断通常比绝对评分更容易且更稳定,这减轻了严格规模校准的需求。验证器也无需自身训练样本。我们将验证器集成到三个已发布的自我进化引擎(GEPA、ADRS、ShinkaEvolve)中,并报告了两种变体:自适应焦点,它保留了引擎现有的验证集父选择;软Elo,它让验证器的判断驱动父选择,从而使验证集奖励也下降。在多个智能体和两种工件基质(提示和代码)上,我们的方法在大多数评估设置中达到或超过了全奖励基线,并且这种模式在跨家族验证器交换中仍然存在。因此,成对门是在具有竞争力的任务准确性下替代每步奖励设计的现成方案,且无需标记成本。

英文摘要

A self-evolving agentic loop repeatedly proposes a tweaked version of an agent (its prompt template or program) and accepts or rejects the change based on a per-iteration quality signal. Designing that signal is often the costly part of the project: a reliable scalar reward requires domain expertise and labeled examples that are themselves as expensive to assemble as the agent's underlying task. We propose replacing the scalar at the accept/reject gate with a pairwise validator: a frozen LLM that, given the parent and child candidate, returns a binary verdict on which is better. Pairwise judgment is generally easier and more stable than absolute scoring, due to its contrastive nature, which mitigates the need for strict scale calibration. The validator also requires no training of its own. We integrate the validator into three published self-evolving engines (GEPA, ADRS, ShinkaEvolve) and report two flavors: Adaptive Focus, which retains the engine's existing val-set parent selection, and Soft Elo, which lets the validator's verdicts drive parent selection so that val-set rewards drop as well. Across multiple agents and two artifact substrates (prompt and code), our method matches or exceeds the full-reward baseline on the majority of settings we evaluate, and the pattern survives a cross-family validator swap. The pairwise gate is thus a drop-in replacement for per-step reward design at competitive task accuracy without the labeling cost.

↑