arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向证书驱动的软件移植:一种用于科学程序优化的自改进智能体框架

Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization

Piyush Jha, Aishik Ghosh, Vijay Ganesh

arXiv 2609.34069首次发表:更新:

发表机构

Georgia Institute of Technology; Lawrence Berkeley National Laboratory(佐治亚理工学院; 劳伦斯伯克利国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出证书驱动进化搜索(CDES),通过记录失败候选的可执行限制来增强LLM进化搜索,应用于Geant4的CPU到GPU移植,实现最高23.54倍加速,并将正确性通过率从55%提升至90%。

AI 中文摘要

大型科学代码库的升级和重写历来是一项重大挑战。虽然基于大型语言模型(LLM)的进化搜索可以移植并加速遗留代码,但仅靠提示中的修复反馈并不能防止后续候选方案重复相同的错误。我们引入了证书驱动进化搜索(CDES),该方法通过从失败候选方案中提取的可执行限制来扩展进化搜索,这些限制被记录为假设证书、检查器证据和合理限制。其控制逻辑通过拒绝、回溯和定向修复来强制执行这些限制,同时保留兼容的编辑。我们将CDES应用于Geant4工具包中两个粒子模拟函数的CPU到GPU翻译,并使用一个超越单元测试的评估框架,该框架结合了形式检查、数值比较、物理检查和GPU安全测试。生成的实现相对于CPU代码实现了13.78倍和23.54倍的函数级加速,包括数据转换和传输;对于其中一个函数,GPU吞吐量超过专家实现14.9%,当组合互补组件时达到16.1%。在针对执行设置的消融实验中,证书反馈将通过所需正确性检查的候选方案比例从55%提高到90%。

英文摘要

The upgrade and rewriting of large scientific codebases has traditionally been a major challenge. While evolutionary search with large language models (LLMs) can port and accelerate legacy code, repair feedback in prompts alone does not prevent subsequent candidates from repeating the same errors. We introduce Certificate-Driven Evolutionary Search (CDES), which extends evolutionary search with enforceable restrictions derived from failed candidates, recorded as certificates of assumptions, checker evidence, and justified restrictions. Its control logic enforces these restrictions through rejection, backtracking, and targeted repair while preserving compatible edits. We apply CDES to CPU-to-GPU translation of two particle-simulation functions from the Geant4 toolkit, evaluated with a harness that goes beyond unit tests to combine formal checks, numerical comparisons, physics checks, and GPU safety tests. Generated implementations achieve 13.78x and 23.54x function-level speedups over CPU code, including data conversion and transfers; for one function, GPU throughput exceeds an expert implementation by 14.9%, reaching 16.1% when complementary components are combined. In an ablation over execution settings, certificate feedback increases the fraction of candidates passing required correctness checks from 55% to 90%.

CommentsSubmitted to ML4PS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑