arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

J-Zero:从零数据开始的挑战者-求解器-评判器协同进化框架

J-Zero: Unified Challenger--Solver--Judge Self-Evolution from Zero Data

Gyouk Chu, Myeongho Jeon, Teresa Yeo, Eunho Yang

arXiv 2608.26582首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

J-Zero是从零数据出发的统一挑战者-求解器-评判器协同进化框架,在可验证与不可验证领域均优于基线,且迭代10次后仍能持续改进。

AI 中文摘要

近年来,自进化语言模型已成为通向超级智能的有前景路径,其优势在于可降低人类监督的成本。尽管在可验证领域已取得显著进展,但不可验证领域的自进化仍探索不足。我们提出了从零数据开始的评判器协同适应框架J-Zero,这是一个支持跨两类领域自我改进的统一挑战者-求解器-评判器协同进化框架。挑战者与求解器通过对抗交互协同进化:挑战者生成难度逐步提升的任务,而求解器学习生成更高质量的任务响应。与此同时,评判器利用偏好对进行协同适应,这些偏好对的排序依据是响应的生成方式(即求解器的答案优于挑战者的答案,且其分解重组后的答案优于其单次生成的答案),而非评判器自身的评分。在可验证领域,J-Zero较基线方法平均提升4.2个百分点;在不可验证领域,平均提升8.0个百分点;且至少经过10次迭代后仍持续改进,而基线方法在2次迭代后便出现性能下降。

英文摘要

Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains less explored. We propose Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge self-evolution framework that supports self-improvement across both domains. The Challenger and Solver co-evolve through an adversarial interaction: the Challenger generates increasingly difficult tasks, while the Solver learns to produce higher-quality responses to them. In parallel, the Judge co-adapts using preference pairs whose ordering is known in advance from how each response was produced, i.e., the Solver's answer over the Challenger's, and the Solver's decomposed-and-recombined answer over its one-shot answer, rather than from the Judge's own scores. J-Zero outperforms the baselines by an average of 4.2 points on verifiable and 8.0 points on unverifiable domains, and continues to improve through at least ten iterations, whereas the baselines degrade after two. Further analysis identifies Judge co-adaptation as the key driver of this sustained improvement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑