arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向用于Solidity智能合约的大语言模型辅助的高质量属性生成

Towards LLM-assisted High-Quality Property Generation for Solidity Smart Contracts

Muhammad Wahid, Shahzaib Khan, Mashhood Ali, Muhammad Hassan, Muhammad Naiman Jalil, Affan Rauf

arXiv 2607.23308首次发表:更新:

AI 中文总结

研究利用大语言模型为Solidity智能合约生成高质量属性,通过变异测试衡量属性质量,评估多种提示技术,发现Gemini Pro 1.5结合提示链平均变异得分高,个别合约表现突出,显示大语言模型在属性生成上有潜力达专家水平。

AI 中文摘要

智能合约的不可变特性使得在其部署到区块链后修复漏洞具有挑战性,这意味着安全漏洞可能在更长时间内面临被利用的风险,因此需要全面的部署前测试。基于属性的测试与模糊测试相结合已被证明是发现漏洞的一种有前途的技术。传统上,系统属性由人类专家编写,耗时且成本高。鉴于大语言模型(LLMs)的最新进展及其理解自然语言和代码语义的能力,有可能生成有效的属性。本研究利用最先进的大语言模型为基于Solidity的智能合约生成高质量属性。我们使用变异测试来衡量生成属性的质量。结果表明,大语言模型有潜力生成接近人类专家编写的高质量属性。我们使用各种提示技术(如零样本、少样本和提示链)对大语言模型进行了广泛评估。总体而言,我们发现Gemini Pro 1.5与提示链相结合时,在所有研究配置中实现了最高平均变异得分25.99%,接近人类编写基准的31.75%。然而,我们对每个合约的分析揭示了显著差异,特别是对于LibBit合约,Gemini Pro 1.5在提示链下实现了74.34%的变异得分,与人类编写的属性(74.83%)相当。这突出表明,虽然平均性能有参考价值,但个别合约层面的结果表明,大语言模型在某些情况下可以匹配专家级别的属性生成。

英文摘要

The immutable nature of smart contracts makes it challenging to fix and patch bugs once they are deployed to a blockchain. This implies that security vulnerabilities may be exposed to possible exploitation for a longer period, necessitating comprehensive pre-deployment testing. Property-based testing combined with fuzzing has proven itself as a promising technique for uncovering vulnerabilities. Traditionally, system properties are written by human experts, which is time-consuming and consequently expensive.With the recent advancement in Large Language Models (LLMs) and their ability to 'understand' natural language and code semantics, it may be possible to generate effective properties. This study, leverages state-of-the-art LLMs to generate high-quality properties for Soliditybased smart contracts. We measure the quality of the generated properties using mutation testing. Our results show that LLMs have the potential to generate high-quality properties that are close to those written by human experts. We extensively evaluate LLMs using various prompting techniques (e.g., zero shot, few shot, and prompt chaining). Overall, we find that Gemini Pro 1.5, when combined with prompt chaining, achieves the highest average mutation score of 25.99% among all studied configurations, closely approaching the human written benchmark of 31.75%. However, our per contract analysis reveals notable variance, particularly for the LibBit contract, where Gemini Pro 1.5 under prompt chaining achieves a mutation score of 74.34%, which is on par with human written properties (74.83%). This highlights that while average performance is informative, individual contract level results demonstrate that LLMs can, in some cases, match expert level property generation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑