arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

强化学习增强的大语言模型智能体用于复杂车辆路径问题

Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems

Yi Chen, Zikang Yu, Jiahai Wang, Jinbiao Chen, Jianpeng Zhou, Zizhen Zhang

arXiv 2609.00859首次发表:更新:

发表机构

School of Computer Science and Engineering, Sun Yat-sen University; National University of Singapore(中山大学计算机科学与工程学院; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对车辆路径问题建模需大量领域知识的局限,提出RLEA多智能体框架,结合强化学习与大语言模型,在48种VRP变体上实现16.67%更高成功率,显著降低运行时错误,验证了方法的有效性。

AI 中文摘要

车辆路径问题(VRP)是一类基础的组合优化问题,广泛应用于各类场景。先进的优化求解器可有效解决这类问题,但为求解器构建复杂VRP变体的模型通常需要大量领域专业知识,这限制了先进优化技术的可及性。本文提出了强化学习增强的大语言模型智能体(RLEA),这是一个旨在自动化构建复杂VRP建模的多智能体框架。RLEA引入了一个通过软Q学习训练的轻量级神经规划器,以高效协调基于大语言模型的智能体的动作。此外,我们为系统配备了进化记忆模块和检索增强生成,使智能体在生成和优化求解VRP的程序时,能同时利用积累的经验和外部求解器知识。我们在各类求解器上评估了48种不同的VRP变体,实验结果表明,RLEA的性能优于之前的最先进方法,实现了16.67%的更高成功率,同时显著降低了运行时错误。这些结果验证了将强化学习与基于大语言模型的推理相结合,对于自动化优化建模非常有效。附录可在以下网址获取:this https URL

英文摘要

Vehicle Routing Problems (VRPs) are fundamental combinatorial optimization problems with widespread applications in various scenarios. The advanced optimization solvers can effectively solve such problems. However, modeling complex VRP variants for solvers often requires substantial domain expertise, which limits the accessibility of advanced optimization technologies. In this paper, we propose Reinforcement Learning Enhanced LLMAgents(RLEA), a multi-agent framework designed to automate the modeling of complex VRPs. RLEA introduces a lightweight neural Planner trained with Soft Q-learning to efficiently orchestrate the actions of LLM-based agents. In addition, we equip the system with an evolutionary memory module and retrieval-augmented generation, enabling the agent to leverage both accumulated experience and external solver knowledge during program generation and refinement for solving VRPs. We evaluated 48 distinct VRP variants across various solvers. The experimental results demonstrate that RLEA outperforms the previous state-of-the-ar method, achieving a 16.67% higher success rate while significantly reducing runtime errors. These results validate that integrating reinforcement learning with LLM-based reasoning is highly effective for automated optimization modeling. The appendix is available at: https://doi.org/10.5281/zenodo.19134435.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑