AI 中文总结
本文提出基于静态依赖分析的自动验证流程,评估OpenAI o3生成的微服务分解的结构一致性,发现零样本与少样本提示的结构一致性相当,且需控制映射覆盖率以避免方法学偏差。
AI 中文摘要
将单体系统分解为微服务是软件现代化的关键活动。尽管大语言模型(LLMs)能从文本需求生成语义合理的分解方案,但目前尚不清楚这些方案是否保留了源代码中实现的结构依赖。本文评估了OpenAI o3为PetClinic和Bookstore系统生成的微服务分解的结构一致性,提出了一种基于静态依赖分析的自动验证流程,并使用依赖保留率(TPD)和依赖违反率(TVD)指标对比了零样本和少样本提示的效果。为控制类到服务映射覆盖率的差异,开展了鲁棒性分析,经标准化后,两种提示策略产生的结构一致性相当,在PetClinic系统上TPD值为68.0%,在Bookstore系统上为83.3%。研究结果表明,对LLM生成的分解方案进行结构评估时,应明确控制映射覆盖率,否则提示策略间的表观差异可能反映方法学偏差,而非真正的架构质量。
英文摘要
Decomposing monolithic systems into microservices is a key activity in software modernization. Although Large Language Models (LLMs) can generate semantically plausible decompositions from textual requirements, it remains unclear whether these proposals preserve the structural dependencies implemented in the source code. This paper evaluates the structural adherence of microservice decompositions generated by OpenAI o3 for the PetClinic and Bookstore systems. We propose an automated validation pipeline based on static dependency analysis and compare zero-shot and few-shot prompting using dependency preservation (TPD) and dependency violation (TVD) metrics. A robustness analysis was conducted to control for differences in class-to-service mapping coverage. After normalization, both prompting strategies produced equivalent structural adherence, achieving TPD values of 68.0% (PetClinic) and 83.3% (Bookstore). The findings demonstrate that structural evaluations of LLM-generated decompositions should explicitly control for mapping coverage, as apparent differences between prompting strategies may otherwise reflect methodological bias rather than genuine architectural quality.