发表机构
Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对偏微分方程求解器代码生成,引入RLVP强化学习训练后框架,通过混合验证器解决可验证性问题,在多PDE族训练单个策略,优于基线且有零样本改进转移,训练后小LLM表现良好,策略有组合性证据。
AI 中文摘要
偏微分方程(PDEs)是科学和工程建模的基础,但构建可靠的数值求解器仍需大量人力,需要离散化方案、稳定性条件和边界处理等方面的专业知识。近期工作开始将PDE求解视为大语言模型(LLMs)的代码生成任务,但现有方法主要在推理时操作。同时,具有可验证奖励的强化学习已成为代码和数学推理的训练后范式,但其验证器通常是二元的。在这项工作中,我们引入了RLVP:具有可验证物理的强化学习,这是一个用于多PDE求解器代码生成的强化学习训练后框架。RLVP通过混合验证器解决了可验证性差距,硬程序有效性检查确保可执行性,而连续物理奖励对函数空间精度和PDE残差一致性进行评分。在跨越双曲、抛物、椭圆和不可压缩流系统的不同PDE族上对单个策略进行训练后,RLVP在PDE基准测试中优于预训练和仅监督的基线,并显示出对未见过的PDE的零样本改进转移。我们表明,用RLVP训练后的较小LLM在分布内PDE求解器生成方面可以优于对前沿模型进行提示。训练后的策略在数值模式中显示出组合性证据:它将从训练中使用的PDE中学到的模板、时间步长方案和边界处理原语重新组合到为未见PDE问题生成的求解器中。
英文摘要
Partial differential equations (PDEs) are foundational to modeling in science and engineering, but constructing reliable numerical solvers remains labor-intensive, demanding expert knowledge of discretization schemes, stability conditions, and boundary treatments. Recent work has begun to frame PDE solving as a code-generation task for large language models (LLMs), yet existing approaches operate primarily at inference time: relying on prompting, debugging, self-refinement, and test-time scaling rather than adapting the model itself. In parallel, reinforcement learning with verifiable rewards has emerged as a post-training paradigm for code and math reasoning, but its verifiers are typically binary: a compiler runs, or a test passes. Such signals discard the graded structure of scientific correctness, where two solvers may both execute and yet differ in solution accuracy by orders of magnitude. In this work, we introduce RLVP: Reinforcement Learning with Verifiable Physics, an RL post-training framework for multi-PDE solver code generation. RLVP addresses this verifiability gap with a hybrid verifier: hard program-validity checks ensure executability, while continuous physics rewards score function-space accuracy and PDE-residual consistency. A single policy is post-trained across diverse PDE families spanning hyperbolic, parabolic, elliptic, and incompressible-flow systems. RLVP improves over both pre-trained and supervised-only baselines on PDE benchmarks, and shows zero-shot improvement transfer to held-out PDEs. We show that a smaller LLM post-trained with RLVP can outperform prompting a frontier model on in-distribution PDE solver generation. The trained policy shows evidence of compositionality in numerical motifs: it recombines stencils, time-stepping schemes, and boundary-handling primitives learned from the PDEs used in training into generated solvers for unseen PDE problems.