追逐即课程,捕获锚定信用:用于零数据大语言模型推理的追逃自博弈
The Chase Is the Curriculum, the Capture Anchors the Credit: Pursuit-Evasion Self-Play for Zero-Data LLM Reasoning
浏览论文内容
中文总结 AI 辅助
本研究提出LURE框架,将零数据自博弈转化为追逃博弈,通过捕获锚定信用机制,在三类推理环境和主干模型上,实现优于先进基线的零数据LLM推理性能。
中文摘要 AI 辅助
带有可验证奖励的强化学习已成为提升大语言模型(LLM)推理能力的主流方法,但它依赖大量人工整理的任务集合。零数据自博弈消除了这种依赖,不过现有方法仅通过探测候选任务来验证可学习性,事后拒绝不符合要求的任务,从未在环境的难度轴上确定任务的放置位置,且仅用稀疏的最终奖励为求解器分配信用。我们将零数据自博弈重新定义为追逃博弈:在LURE框架中,LLM逃逸体在每个环境的难度轴上调整任务位置,以领先于规划器-执行器追捕体一步,该追捕体通过可验证交互追捕逃逸体。逃逸体在捕获前沿奖励的驱动下进行训练,当求解器在其恰好一半的回合中捕获它时,该奖励达到峰值,将勉强可捕获转变为学习到的定位策略,而非手动调整的拒绝带。追捕体获得由捕获锚定的密集过程信用,其中单调验证器进度与最终捕获结果在基于回合的KL散度下进行联合分组归一化,以保持协同进化的稳定性。在三个可验证推理环境和三类主干模型上,LURE在统一/专用设置下均优于先进基线,而统一模型在三类任务族的九个保留基准上,比所有训练后的基线实现更强的分布外零样本综合准确率。
英文摘要
Reinforcement learning with verifiable rewards has become the dominant recipe for improving large language model reasoning, yet it presumes large human-curated task collections. Zero-data self-play removes this dependency, but existing methods vet learnability only by probing candidates and rejecting post hoc, never learning where along an environment's difficulty axis to place a task, and credit the solver with sparse terminal rewards alone. We recast zero-data self-play as a pursuit-evasion game: in LURE, an LLM evader positions tasks along each environment's difficulty axis to stay one step ahead of a planner-executor pursuer that hunts it down through verifiable interaction. The evader is trained on a capture-frontier reward that peaks when the solver captures it on exactly half of its rollouts, turning barely catchable into a learned positioning strategy rather than a hand-tuned rejection band. The pursuer earns capture-anchored dense process credit, in which monotone verifier progress is group-normalized jointly with the terminal capture under a round-anchored KL that keeps the co-evolution stable. Across three verifiable reasoning environments and three backbone families, LURE outperforms advanced baselines under unified/specialist settings, while the unified model attains stronger aggregate OOD zero-shot accuracy than all trained baselines across nine held-out benchmarks from three task families.