面向验证的25万行以上遗留气象模拟代码的AI辅助GPU移植
Validation-Centric AI-Assisted GPU Porting of a 250,000+ Line Legacy Weather Simulation Code
浏览论文内容
中文总结 AI 辅助
本文以遗留Fortran气象模拟代码CReSS为例,提出面向验证的AI辅助GPU移植工作流,为162个内核生成经数值验证的GPU实现,获5.1倍应用级加速,还检测到5个内核的数值不一致。
中文摘要 AI 辅助
大型语言模型的最新进展使基于命令行界面(CLI)的智能体成为加速大型遗留科学应用GPU移植的实用工具。然而,这类应用并非仅是旧代码库,而是通过长期开发、与观测结果对比及在领域研究中应用积累了可信度的科学资产。因此,GPU移植必须在适配以GPU为中心的高性能计算(HPC)系统的同时,保留这种科学有效性。本文以CReSS为例,提出了一种面向验证的AI辅助GPU移植工作流;CReSS是一款拥有超过25万行代码的遗留Fortran气象模拟代码。该工作流利用AI智能体提取OpenMP区域,从具有物理意义的模拟状态生成基于转储的内核基准,应用OpenACC转换,并通过与转储的参考数据进行逐元素比较以及应用级验证来验证结果。通过一次真实的台风模拟,该工作流为162个目标内核生成了经数值验证的GPU实现,并在实际的时钟开发成本内实现了5.1倍的应用级加速。特别值得注意的是,它检测到5个内核中因浮点和内在函数差异导致的数值不一致,包括阈值敏感的分支发散和抵消效应,从而能够向应用开发者提供反馈。该案例研究表明,对于需要基于转储验证的大型遗留科学应用,实用的AI辅助GPU移植必须管理跨会话的上下文、运行时状态重建,以及从静态分析的小疏漏中进行代价高昂的恢复。这些发现表明,AI辅助GPU移植不仅需要代码生成,还需要面向验证的工作流设计。
英文摘要
Recent advances in large language models have made CLI-based AI agents a practical tool for accelerating GPU porting of large legacy scientific applications. Such applications, however, are not merely old code bases; they are scientific assets whose credibility has been accumulated through long-term development, comparison with observations, and use in domain studies. GPU porting must therefore preserve this scientific validity while adapting the implementation to GPU-centric HPC systems. This paper presents a validation-centric AI-assisted GPU porting workflow through a case study of CReSS, a legacy Fortran weather simulation code with more than 250,000 lines. The workflow uses an AI agent to extract OpenMP regions, generate dump-based kernel benchmarks from physically meaningful simulation states, apply OpenACC transformations, and validate results through element-wise comparison with dumped reference data and application-level validation. Using a real typhoon simulation, the workflow produced numerically validated GPU implementations for 162 target kernels and achieved a 5.1x application-level speedup within practical wall-clock development cost. In particular, it detected numerical discrepancies in five kernels caused by floating-point and intrinsic-function differences, including threshold-sensitive branch divergence and cancellation effects, enabling feedback to the application developers. The case study suggests that, for large legacy scientific applications requiring dump-based validation, practical AI-assisted GPU porting must manage session-spanning context, runtime-state reconstruction, and costly recovery from small static-analysis omissions. These findings demonstrate that AI-assisted GPU porting requires not only code generation, but validation-centric workflow design.