arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过求解器反馈教会大语言模型生成具有挑战性的MILP实例

Teaching LLMs to Generate Challenging MILP Instances via Solver Feedback

Jitin Singla, Parikshit Pareek, Pratik Jawanpuria, Parag Singla

arXiv 2609.37356首次发表:更新:

发表机构

IIT Roorkee; IIT Bombay; IIT Delhi(印度理工学院罗克分校; 印度理工学院孟买分校; 印度理工学院德里分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出利用求解器反馈设计奖励,通过挑战者-求解器自博弈微调大语言模型,生成可行且高难度的MILP实例,显著提升求解器搜索节点和间隙。

AI 中文摘要

生成既可行又具有计算挑战性的优化实例,对于基准测试求解器和训练基于学习的优化算法至关重要。现有的非大语言模型生成器依赖于种子实例或参数调整,导致测试时计算成本高昂,而现有的大语言模型生成器缺乏显式的难度度量。最近的带验证器反馈的强化学习方法仅评估二元正确性,这与生成具有挑战性的问题不一致。我们注意到,优化求解器在其流程的多个阶段报告求解成本,并利用这一点设计了一个奖励函数,该函数同时评分生成问题的可解性和难度,难度通过分支定界节点数和切割后松弛间隙来衡量。我们的关键思想是一种挑战者-求解器不对称自博弈方法,其中大语言模型挑战者生成逐渐更难的实例,求解器验证可行性和难度,因此不需要种子或训练MILP实例。我们使用GRPO和规模课程微调Gemma-4-12B和Qwen3.5-4B,得到OptiScribe-12B和OptiScribe-4B,它们能从自然语言指令生成可行且具有挑战性的MILP问题。在容量设施选址和最大割问题上,OptiScribe-12B将SCIP搜索节点中位数提高了1.7-5倍,切割后间隙提高了1.1-1.7倍(相对于其基础模型),并将设施选址的可行性率提高了9-19个百分点,而OptiScribe-4B将节点中位数提高了最多15.6倍。生成的问题覆盖了比同规模公共基准更广的难度范围,遵循关于密度和难度的指令,并能调整公共库缺乏的求解器设置家族。这些结果表明,在自博弈模式中使用的优化特定奖励,可以教会大语言模型生成高难度的优化基准。我们将在论文被接受后公开发布代码和模型。

英文摘要

Generating optimization instances that are both feasible and computationally challenging is crucial for benchmarking solvers and training learning-based optimization algorithms. Existing non-LLM generators rely on seed instances or parameter tuning, resulting in high test-time computational cost, while existing LLM generators lack explicit hardness measures. Recent reinforcement learning methods with verifier feedback evaluate only binary correctness, which is misaligned with generating challenging problems. We note that an optimization solver reports the cost of solving at several stages of its pipeline, and leverage this to design a reward that scores both the solvability and the hardness of generated problems, measured by branch-and-bound nodes and post-cut relaxation gaps. Our key idea is a challenger-solver asymmetric self-play approach, where an LLM challenger generates progressively harder instances and the solver verifies feasibility and hardness, so no seed or training MILP instances are required. We fine-tune Gemma-4-12B and Qwen3.5-4B with GRPO and a size curriculum into OptiScribe-12B and OptiScribe-4B, which generate feasible yet challenging MILP problems from natural language instructions. On capacitated facility location and max-cut, OptiScribe-12B raises median SCIP search nodes by 1.7-5x and post-cut gaps by 1.1-1.7x over its base model and improves the feasibility rate on facility location by 9-19 points, while OptiScribe-4B raises median nodes by up to 15.6x. The problems cover a wider difficulty range than public benchmarks of the same size, follow instructions on density and difficulty, and can tune solver settings for families that public libraries lack. These results indicate that optimization-specific rewards, used in self-play mode, can teach LLMs to generate high-difficulty optimization benchmarks. We will release our code and models publicly on acceptance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑