arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过求解器驱动的自我对弈学习优化

Learning to Optimize through Solver-Grounded Self-Play

Xia Jiang, Yaoxin Wu, Chenyu Zhou, Mengzhu Xu, Wim P. M. Nuijten, Yingqian Zhang

arXiv 2609.34205首次发表:更新:

发表机构

Eindhoven University of Technology; Shanghai Jiaotong University(埃因霍温理工大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出OPT-Zero,首个无需外部数据的自我对弈训练框架,通过提议者与求解者的双角色闭环及求解器反馈强化学习,使LLM在优化建模中达到数据依赖方法水平并显著提升泛化能力。

AI 中文摘要

优化建模是许多决策场景的核心,但传统上需要广泛的专业知识。尽管大型语言模型(LLMs)在自动化这一过程中展现出潜力,当前的训练范式主要依赖于人工标注或教师生成的数据集。这种依赖引入了泛化天花板(Generalization Ceiling),即模型过度拟合于狭窄的数据分布,以及能力锚定(Capability Anchoring),即模型的推理受限于标注者的熟练程度和教师模型的能力。针对这一问题,我们提出了OPT-Zero,这是首个完全自我对弈的优化建模训练框架,无需任何外部训练数据。OPT-Zero在双角色闭环中使用单个LLM:一个提议者(Proposer),负责综合生成日益具有挑战性的优化问题及其数学公式和求解代码;以及一个求解者(Solver),仅根据自然语言的问题描述尝试解决这些问题。基于外部优化求解器的执行反馈,我们使用强化学习交替训练这两个角色。这一过程促进了自动课程学习,其中提议者和求解者共同进化:提议者生成更难的有效问题,无缝地增强了求解者的结构推理能力。大量结果表明,在零策划数据的情况下,OPT-Zero匹配了最先进的数据依赖方法,同时展现出显著更强的泛化能力,确立了自我对弈训练作为推进LLM在优化问题建模和求解中推理能力的高度可扩展范式。

英文摘要

Optimization modeling is central to many decision-making scenarios, but traditionally requires extensive domain expertise. While Large Language Models (LLMs) have shown promise in automating this process, current training paradigms mainly rely on human-annotated or teacher-generated datasets. This dependence introduces a Generalization Ceiling, where models overfit to narrow data distributions, and Capability Anchoring, where models' reasoning is bounded by annotator proficiency and teacher model capability. In response, we propose OPT-Zero, the first fully self-play training framework for optimization modeling that requires zero external training data. OPT-Zero employs a single LLM in a dual-role closed loop: a Proposer that synthesizes increasingly challenging optimization problems alongside their mathematical formulations and solving code, and a Solver that attempts to resolve the problems given only natural-language problem descriptions. Grounded in execution feedback from external optimization solvers, we alternately train both roles using reinforcement learning. This process fosters an auto-curriculum in which the Proposer and Solver co-evolve: generating harder valid problems by the Proposer seamlessly enhances the structural reasoning ability of the Solver. Extensive results indicate that with zero curated data, OPT-Zero matches state-of-the-art data-dependent methods while exhibiting substantially stronger generalizability, establishing self-play training as a highly scalable paradigm for advancing LLM reasoning in modeling and solving optimization problems.

Comments36 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑