发表机构
Purdue University; LightSpeed Studios(普渡大学; 光速工作室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出M$^3$P-R1,利用强化学习微调大语言模型,将多模态运动规划任务分解为MIP代码生成,实现稳健可验证的求解。
AI 中文摘要
多模态运动规划(M$^3$P)需要对连续运动和离散模式转换进行联合推理,这使得高效求解变得困难。例如,一个双足机器人可能先走到目标位置,然后用其手臂抓取物体。这一场景同时涉及模式转换和连续动力学,产生的可行路径是纯离散或纯连续规划器都无法处理的。虽然混合整数规划(MIP)提供了一个原则性框架,但为非凸问题构建易处理的公式通常是手动的且特定于领域,尤其是在非凸机器人任务所需的基于近似、离散化的MIP体系中。我们提出了M$^3$P-R1,一种强化学习方法,通过微调大语言模型(LLMs)将M$^3$P任务分解为MIP变量、约束和目标。该模型不是直接输出答案(这往往容易出现幻觉),而是生成使用MIP优化库和约束接口的可执行Python代码。这使得基于求解器的执行能够提供稳健且可验证的解决方案。通过针对求解器的基于结果的奖励进行训练,M$^3$P-R1学会了组合模态级离散化原语并合成跨模态耦合约束,从而为复杂的M$^3$P任务生成可执行的MIP程序。
英文摘要
Multi-Modal Motion Planning (M$^3$P) requires joint reasoning over continuous motions and discrete mode transitions, making it difficult to solve efficiently. For instance, a bipedal robot may walk to a target location and then use its arms to grasp an object. This scenario captures both mode transitions and continuous dynamics, yielding feasible paths that neither purely discrete nor continuous planners can handle. While Mixed-Integer Programming (MIP) offers a principled framework, constructing tractable formulations for non-convex problems is typically manual and domain-specific, especially in the approximate, discretization-based MIP regime needed for non-convex robotic tasks. We propose M$^3$P-R1, a reinforcement learning method that fine-tunes large language models (LLMs) to decompose M$^3$P tasks into MIP variables, constraints, and objectives. Instead of directly outputting answers, which are often prone to hallucination, the model generates executable Python code using MIP optimization libraries and constraint interfaces. This enables solver-backed execution for robust and verifiable solutions. Trained with an outcome-driven reward against the solver, M$^3$P-R1 learns to compose modality-level discretization primitives and synthesize cross-modal coupling constraints, producing executable MIP programs for complex M$^3$P tasks.
Comments56 pages, 22 figures