发表机构
Leiden University; NWO-I ASTRON(莱顿大学; 荷兰科学研究组织下属射电天文学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究利用小型语言模型(SLMs)优化射电天文学代码,通过多采样生成策略和纳入编译器反馈两种方式增强SLMs,提高代码生成质量和性能,用更少计算资源匹配或超越大型单生成模型,且使所有测试模型持续改进。
AI 中文摘要
近期的大型语言模型(LLMs)能够生成并优化复杂代码。我们研究利用LLMs为大规模科学领域生成并优化代码,重点关注射电天文学与可持续性。LOFAR望远镜正在升级,观测天空区域大幅增加且数据处理更快,但计算需求预计增长40倍,这依赖于现有软件的严格性能优化及加速器的广泛应用,而代码库庞大任务艰巨。我们研究并展示一种人工智能驱动方法,辅助开发者评估和优化代码,包括移植到硬件加速器。LOFAR社区致力于可持续解决方案,需在不增加能源预算的情况下实现改进。因LLMs能耗大,我们提议用小型语言模型(SLMs)以限制环境影响。本文展示如何通过智能人工智能增强SLMs,从多采样生成策略和纳入编译器反馈两方面扩展SLMs以提高代码生成质量和性能。证明多采样SLMs能用更少计算资源匹配或超越大型单生成模型,且将编译器输出反馈到SLMs能使所有测试模型持续改进。我们的方法通用,还可在代码生成管道中使用检索增强生成(RAG)以及静态和动态分析工具。
英文摘要
Recent Large Language Models (LLMs) can produce and optimize complex code. We investigate the use of LLMs to generate and optimize code for large-scale sciences, focusing on radio astronomy and sustainability. The LOFAR telescope is currently being upgraded, significantly increasing the sky area observed, while simultaneously processing more data faster. However, this is expected to increase the computational requirements 40-fold. This upgrade thus critically depends on rigorous performance optimization of existing software and widespread adoption of accelerators. The code base is very large, making this a daunting task. We therefore investigate and demonstrate an AI-driven approach meant to assist developers in evaluating and optimizing their code, including porting to hardware accelerators. The LOFAR community is committed to sustainable solutions, and needs to achieve these improvements without increasing the energy budget. We thus need to optimize existing codes or port them to accelerators, while making sure that the optimization process itself is also energy efficient. This poses a challenge, since LLMs are energy-intensive. We therefore propose to use Small Language Models (SLMs) instead to limit environmental impact. In this paper, we show how to enhance SLMs through the use of agentic AI. We extend the SLMs in two ways to improve code generation quality and performance: first with a multi-sampling generation strategy and second with incorporating compiler feedback. We demonstrate that multi-sampling SLMs can match or surpass larger single-generation models with fewer computational resources and that feeding compiler output back into the SLMs leads to consistent improvements across all tested models. Our approach is generic, and can also use Retrieval Augmented Generation (RAG) as well as static and dynamic analysis tools in the code generation pipeline.