发表机构
Federal University of Campina Grande (UFCG); VIRTUS/UFCG; Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET); Universidad Nacional del Centro de la Provincia de Buenos Aires (UNICEN); Federal Institute of Pernambuco (IFPE); Rui Barbosa State School(帕伊巴联邦大学(UFCG); VIRTUS/UFCG; 国家科学技术研究委员会(CONICET); 布宜诺斯艾利斯省中部国立大学(UNICEN); 伯南布哥联邦学院(IFPE); 鲁伊·巴尔博萨州立学校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨OpenAI o3能否仅从文本需求生成微服务架构,经实验发现少样本提示下其服务识别与通信恢复的一致性优于零样本,为需求驱动的架构合成提供了潜力。
AI 中文摘要
微服务架构已成为单体系统现代化的主流方式,但识别合适的服务仍具挑战性,且在很大程度上依赖人工。现有分解方法主要以代码为中心,限制了其在仅存在文本需求的早期设计阶段的适用性。尽管大语言模型(LLM)取得了进展,但关于其从自然语言需求(包括服务定义和服务间交互)合成完整微服务架构能力的实证证据有限。本研究探讨LLM是否能衔接需求工程与架构设计,仅从文本需求生成架构,并评估结果的结构一致性和感知质量。我们使用OpenAI o3开展混合方法研究,在零样本(ZS)和少样本(FS)提示下针对两个系统(Bookstore、PetClinic)各执行一次。架构通过以下方式评估:(i)与参考架构对比,计算服务识别和通信恢复的精确率、召回率和F1值;(ii)对正确性、完整性、模块化和合理性进行盲法专家评估,并综合开放式反馈。OpenAI o3在少样本提示下的服务识别一致性更高(零样本F1值为0.79,少样本为0.97)。通信恢复更具挑战性:零样本生成的架构密集,召回率高但精确率低(F1值为0.61),而少样本提升了一致性,F1值达0.82且减少了无支持的依赖关系。专家评估证实了这些结果,少样本生成的架构被认为比零样本输出更具模块化、连贯性和合理性。OpenAI o3在示例提示引导下展现出需求驱动合成的潜力,但结果来自两个小型系统,具有模型和上下文特异性,并非与模型无关的证明。
英文摘要
Microservice architectures have become dominant for modernizing monolithic systems, yet identifying appropriate services remains challenging and largely manual. Existing decomposition approaches are predominantly code-centric, limiting applicability in early design stages where only textual requirements are available. Despite advances in Large Language Models (LLMs), limited empirical evidence exists on their ability to synthesize complete microservice architectures from natural-language requirements, including service definitions and inter-service interactions. This study investigates whether an LLM can bridge requirements engineering and architectural design, generating architectures solely from textual requirements and evaluating structural agreement and perceived quality of results. We conduct a mixed-method study using OpenAI o3 under zero-shot (ZS) and few-shot (FS) prompting across two systems (Bookstore, PetClinic), one execution per system/condition. Architectures are evaluated through (i) comparison with reference architectures using precision, recall, and F1-score for service identification and communication recovery, and (ii) a blinded expert assessment of correctness, completeness, modularity, and plausibility, plus open feedback synthesis. OpenAI o3 identifies services with higher agreement under FS prompting (F1 = 0.79 for ZS versus = 0.97 for FS). Communication recovery is more challenging: ZS produces dense architectures with high recall but low precision (F1 = 0.61), while FS improves agreement, reaching F1 = 0.82 and reducing unsupported dependencies. Expert evaluation corroborates these results, with FS architectures perceived as more modular, coherent, and plausible than ZS outputs. OpenAI o3 shows potential for requirements-driven synthesis when guided by exemplar prompting. Results are model- and context-specific from two small systems, not model-independent proof.