arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReqGenX:对遗留SRS文档的原子分解、工件再生和重建的实证研究

ReqGenX: An Empirical Study of Atomic Decomposition, Artifact Regeneration, and Reconstruction for Legacy SRS Documents

Ragib Shahariar Ayon, Rayed Fahmi, Sumon Biswas, Shibbir Ahmed

arXiv 2607.16564首次发表:更新:

发表机构

Texas State University; Case Western Reserve University(德克萨斯州立大学; 凯斯西储大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在将遗留SRS文档转化为可追溯工件,以细粒度评估基于大语言模型的SRS生成。方法是用ReqGenX进行实证研究,包括分解、投票、生成工件等步骤。结果表明该方法有效,能支持评估并揭示相关权衡。

AI 中文摘要

背景:评估自动化软件需求规范(SRS)生成具有挑战性,因为很少有数据集能提供源需求、中间引出工件和生成规范之间的细粒度可追溯性。目标:研究遗留SRS文档能否转化为可追溯的合成预SRS工件,以支持对基于大语言模型的SRS生成进行细粒度评估。方法:使用ReqGenX进行实证研究,它将SRS部分分解为基于源的原子语句,通过多语言模型多数投票将原子路由到受标准启发的工件类型,并使用带迭代判断引导细化的受限提示生成工件。我们使用七种PURE SRS文档,通过基础、质量、信息保留和下游重建分析来评估ReqGenX。结果:ReqGenX产生忠实且可用的原子,对齐分数中位数通常在0.96至0.99之间,Prometheus分数在4.34至4.85之间。生成的工件在其源原子中保持高度基础,对齐分数值通常在0.80 - 0.94之间,判断通过率接近100%;更严格的Prometheus评估产生的通过率从54.8%到97.1%。在下游SRS重建案例研究中,工件支持的原子可从生成的SRS中恢复,SBERT均值在0.69至0.75之间,对齐分数中位数在0.76至0.84之间。结论:可追溯的合成预SRS工件可以支持对基于大语言模型的SRS生成进行更细粒度的评估,同时揭示忠实性、信息保留和工件完整性之间的权衡。

英文摘要

Background: Evaluating automated Software Requirements Specification (SRS) generation is challenging because few datasets provide fine-grained traceability between source requirements, intermediate elicitation artifacts, and generated specifications. Aims: We aim to study whether legacy SRS documents can be transformed into traceable synthetic pre-SRS artifacts that support fine-grained evaluation of LLM-based SRS generation. Method: We conduct an empirical study using ReqGenX, a controlled pipeline that decomposes SRS sections into source-grounded atomic statements, routes atoms to standards-inspired artifact types through multi-LLM plurality voting, and generates artifacts using constrained prompts with iterative judge-guided refinement. We evaluate ReqGenX on seven PURE SRS documents using grounding, quality, information retention, and downstream reconstruction analyses. Results: ReqGenX produces faithful and usable atoms, with median AlignScore values typically between 0.96 and 0.99 and Prometheus scores ranging from 4.34 to 4.85. Generated artifacts remain strongly grounded in their source atoms, with AlignScore values typically between 0.80--0.94 and judge pass rates near 100%; stricter Prometheus evaluation yields pass rates from 54.8% to 97.1%. In a downstream SRS reconstruction case study, artifact-backed atoms remain recoverable from generated SRSs, with SBERT means between 0.69 and 0.75 and AlignScore medians between 0.76 and 0.84. Conclusions: Traceable synthetic pre-SRS artifacts can support more fine-grained evaluation of LLM-based SRS generation, while exposing tradeoffs among faithfulness, information retention, and artifact completeness.

CommentsAccepted at the International Symposium on Empirical Software Engineering and Measurement (ESEM) technical papers (main) track 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑