发表机构
School of Artificial Intelligence, University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences(中国科学院大学人工智能学院; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对现有PDE求解器合成方法中自回归模型解码冗余的问题,提出基于离散扩散语言模型的DiffPDE框架,结合局部重掩码填充策略与ID-GRPO强化学习方案,在PDEBench上实现了更优准确率与更快修复速度。
AI 中文摘要
现有的合成偏微分方程(PDE)求解器的方法主要依赖自回归模型,然而在处理本质上局部化的错误时,其全局从左到右的解码会产生大量冗余。在本研究中,我们挑战这一低效范式并提出DiffPDE,一个利用离散扩散语言模型进行针对性代码修复的框架。通过引入局部重掩码与填充策略,DiffPDE仅重新生成错误区域,同时保留正确上下文,使生成过程自然契合PDE错误的稀疏特性。此外,为处理需要顺序干预的耦合错误,我们提出迭代调试GRPO(ID-GRPO),这是一种强化学习方案,可通过中间奖励在单条轨迹内实现多轮调试。在PDEBench上的实验表明,DiffPDE达到了有竞争力的准确率,优于同等规模的自回归(AR)模型,且显著加快了修复速度。
英文摘要
Existing approaches for synthesizing Partial Differential Equation (PDE) solvers predominantly rely on autoregressive models, yet their global left-to-right decoding incurs substantial redundancy when addressing inherently localized bugs. In this work, we challenge this inefficient paradigm and propose DiffPDE, a framework leveraging discrete diffusion language models for targeted code repair. By introducing a localized re-masking and infilling strategy, DiffPDE regenerates only erroneous regions while preserving correct context, naturally aligning generation with the sparse nature of PDE errors. Furthermore, to handle coupled bugs requiring sequential interventions, we present Iterative Debugging GRPO (ID-GRPO), a reinforcement learning scheme that enables multi-round debugging within single trajectories via intermediate rewards. Experiments on PDEBench show that DiffPDE achieves competitive accuracy, outperforms same-scale AR models, and significantly accelerates repair.