AI 中文总结
本研究提出差分故障注入验证方法,以GAMESS为对象测试LLM现代化科学软件的正确性,发现转换后积分核与原始核一致,还暴露了精度降低时的并行死锁等问题。
AI 中文摘要
大语言模型(LLM)智能体正越来越多地用于现代化支撑生产级科学软件的遗留Fortran代码,但这些转换的验证仅强调标称执行,可能未测试现代化是否保留了原始代码对故障、扰动和精度降低的响应。我们提出一种差分故障注入验证方法:该工具在12个位点对GAMESS的共享自洽场驱动程序进行插桩,向原始实现和LLM现代化实现施加相同的确定性故障,隔离转换后的积分核。在超过2200次运行中,瞬态故障吸收成本与基于收缩的模型匹配(预测斜率为每比特0.74和1.49次迭代,实测为0.82和1.50次);持续扰动使每额外比特的最终能量误差减半;该测试还暴露了精度降低时与阶段相关的并行死锁和假收敛。原始核与现代化核在全部200次配对注入中均一致,且测量引导的同步变更与现代化结合后,在全部40次配对中均匹配。
英文摘要
Large language model (LLM) agents are increasingly used to modernize the legacy Fortran underlying production scientific software, but validation of these transformations emphasizes nominal executions and may not test whether a modernization preserves the original code's response to faults, perturbations, and reduced precision. We present a differential fault-injection validation method: a harness instruments the shared self-consistent-field driver of GAMESS at twelve sites and applies identical, deterministic faults to the original and LLM-modernized implementations, isolating the converted integral kernels. Across more than 2,200 runs, transient-fault absorption costs match a contraction-based model (predicted slopes 0.74 and 1.49 iterations per bit; measured 0.82 and 1.50), persistent perturbations halve final-energy error per additional bit, and the campaigns expose phase-dependent parallel deadlocks and false convergence under reduced precision. The original and modernized kernels agree in all 200 paired injections, and a measurement-guided synchronization change composes with the modernization, matching in all 40 pairs.