arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34427cs.LGcs.CL

LLMs作为自适应元求解器:面向工业规模优化的策略多样化强化学习

LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization

  • Tokentide AI
  • Alibaba Group(阿里巴巴集团)
  • East China Normal University(华东师范大学)
  • Fudan University(复旦大学)
  • The University of Hong Kong(香港大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Shihao Zhang, Weiting Liu, Siyu Shao, Yitian Chen, Jianfeng Feng, Dongdong Ge, Yinyu Ye

AI总结:

提出策略多样化强化学习(SDRL)框架,利用求解器集成推理、精确算法与启发式搜索的互补性,训练LLM作为自适应元求解器,在工业规模优化任务上超越现有微调方法和前沿模型。

AI中文摘要:

将基于LLM的优化从教科书规模的实例扩展到现实世界的工业任务仍然是一个关键且未解决的挑战。现有方法主要在小规模、自包含的文本问题上进行评估,并且通常采用求解器集成范式,这限制了它们处理实际优化工作负载的规模和结构多样性的能力。在这项工作中,我们提出了一个实用框架,用于训练开源LLM以解决现实世界中的工业规模优化问题。我们首先通过实验证明,求解器集成推理、精确组合算法和启发式搜索在不同的问题结构和规模上展现出互补的优势。受此启发,我们引入了策略多样化强化学习(SDRL),它将LLM训练为自适应优化元求解器。SDRL通过一种基于正确性门控的层次化多样性奖励来利用这种互补性,该奖励促进在不同策略内部及策略之间进行稳健的探索,有效防止策略过早崩溃。我们还引入了一种混合格式训练方案,该方案同时支持自包含的文本问题和基于文件的实例。在全面的评估中,我们的框架在基准测试的平均表现以及工业规模优化任务上均优于现有的微调方法和前沿模型,包括DeepSeek-V4-Pro和GPT-5.5。

英文摘要:

Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated paradigm, limiting their ability to handle the scale and structural diversity of practical optimization workloads. In this work, we propose a practical framework for training open-source LLMs to tackle real-world, industrial-scale optimization. We first show empirically that solver-integrated reasoning, exact combinatorial algorithm, and heuristic search exhibit complementary strengths across different problem structures and scales. Motivated by this, we introduce Strategy-Diverse Reinforcement Learning (SDRL), which trains LLMs as adaptive optimization meta-solvers. SDRL leverages this complementarity through a correctness-gated hierarchical diversity reward that promotes robust exploration across varying strategies and within each strategy, effectively preventing premature strategy collapse. We further introduce a mixed-format training scheme that jointly supports both self-contained textual problems and file-grounded instances. Across comprehensive evaluations, our framework outperforms existing fine-tuned methods and frontier models including DeepSeek-V4-Pro and GPT-5.5, both on average across benchmarks and on industrial-scale optimization tasks.

↑