AI 中文总结
本文通过对比五种最先进LLMs与GraphWalker工具及其内置算法,针对四类不同复杂度应用模型开展实证评估,探究LLMs用于自动化基于模型的测试生成的效果,发现其可优化缩短测试路径与步长。
AI 中文摘要
大型语言模型在软件工程任务尤其是软件测试中展现出强大潜力。基于模型的测试(MBT)是一种软件测试技术,为解决MBT在工业应用中面临的广泛可扩展性挑战,本文开展了一项针对大型语言模型(LLMs)用于自动化基于模型的测试生成的实证评估,将其与最先进的基于模型的测试工具GraphWalker及其内置算法(边和顶点覆盖设置下的随机算法与快速随机算法)进行对比。评估针对复杂度逐步提升的四个GraphWalker模型(两个Web应用Parabank和Testinium,以及两个硬件应用TLC和RISC-V),结果显示,使用最新的五个最先进LLMs(GPT-5.1、GPT-5.2、Claude Opus 4.5、Claude Sonnet 4.5及Gemini 2.5 Pro)可有效优化并缩短测试路径与步长,展现出巨大潜力。
英文摘要
Large language models have shown strong potential for software engineering tasks, particularly software testing. Model-based testing (MBT) is a software testing technique. To address the broad scalability challenge for industrial adoption of MBTs, our paper presents an empirical evaluation of Large Language Models (LLMs) for automated model-based test generation, compared with a state-of-the-art model-based testing tool (GraphWalker) and its built-in algorithms (random and quick random for edge and vertex coverage settings). Our evaluation indicates strong potential to optimize and shorten test paths and step sizes using the recent five state-of-the-art LLMs (GPT-5.1, GPT-5.2, Claude Opus 4.5, Claude Sonnet 4.5, and Gemini 2.5 Pro) against four GraphWalker models (two web applications (Parabank and Testinium) and two hardware applications (TLC and RISC-V) ) of escalating complexity.
CommentsAccepted at UYMS 2026 (17th National Software Engineering Symposium), Mugla, Turkiye, May 14-16, 2026