发表机构
University of Helsinki; Free University of Bozen-Bolzano(赫尔辛基大学; 博尔扎诺自由大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过系统文献综述分析20项研究,发现LLM主要用作微服务黑盒模糊测试的语义输入生成器,可提升输入有效性与测试效果,但存在评估异质性等挑战,为相关设计部署提供支撑。
AI 中文摘要
微服务系统(MSS)日益依赖异构应用程序编程接口(API),其组合输入空间与有状态依赖关系对传统模糊测试构成挑战。与此同时,大语言模型(LLM)近期被引入,通过对规范、输入及运行时反馈的语义推理增强模糊测试。本文针对微服务的LLM辅助模糊测试开展系统文献综述(SLR),以综合LLM的应用方式、评估情况及尚存挑战。遵循既定SLR指南,我们分析了2024至2026年间发表的20项主要研究。结果显示,LLM主要用作黑盒模糊测试中的语义输入生成器,且正逐渐转向基于智能体(agent)和检索增强架构,提升了有效输入生成能力,适度提高了代码覆盖率与漏洞检测率。不过,评估仍存在异质性,基准标准化不足、成本报告稀缺,且偏向单服务实验,凸显出与现实多服务系统的差距。本综述提供了LLM角色与集成方式的分类、评估实践的统一视角,以及开放挑战与研究方向的映射,为微服务中LLM驱动模糊测试的设计与部署提供支持。
英文摘要
Microservice systems (MSS) increasingly rely on heterogeneous APIs whose combinatorial input space and stateful dependencies challenge traditional fuzz testing. Meanwhile, Large Language Models (LLMs) have recently been introduced to enhance fuzzing with semantic reasoning over specifications, inputs, and runtime feedback. This paper presents a systematic literature review (SLR) of LLM-assisted fuzz testing for microservices to synthesise how LLMs are applied, evaluated, and what challenges remain. Following established SLR guidelines, we analyze 20 primary studies published between 2024 and 2026. Results show LLMs are mainly used as semantic input generators in black-box fuzzing, with a growing shift towards agent-based and retrieval-augmented architectures, improving valid input generation and modestly increasing coverage and vulnerability detection. However, evaluation remains heterogeneous, with limited benchmark standardization, scarce cost reporting, and a bias toward single-service experiments, highlighting a gap with real-world multi-service systems. This review provides a taxonomy of LLM roles and integrations, a consolidated view of evaluation practices, and a mapping of open challenges to research directions, supporting the design and deployment of LLM-driven fuzzing in microservices.