用于需求工程的大语言模型:跨任务实证评估
Large Language Models for Requirements Engineering: A Cross-Task Empirical Evaluation
AI总结:
本文通过两项互补实证研究,开展首个涵盖五项需求工程相关活动的大语言模型跨任务实证评估,明确其性能高度依赖任务,为LLMs在RE中的实际应用提供复现材料与能力认知。
AI中文摘要:
需求相关信息分散在用户反馈、开发者讨论、软件仓库等异构工件中,这使得提取可操作的需求知识既耗费人力又难以规模化。大语言模型(LLMs)可支持诸多需求工程(RE)活动,涵盖分类、可追溯性识别到规格说明与解释生成,但现有证据在任务、工件类型和评估设置间分散,且研究很少提供跨任务评估或复现包。本文开展两项互补的实证研究,在五项RE相关活动中评估LLMs:第一项是针对五个轻量级开源LLMs的受控实验,用于反馈驱动的需求分类和规格说明生成;第二项是针对两个前沿LLMs的探索性工业案例研究,使用真实项目工件开展可追溯性链接识别和可追溯性解释生成。分类和可追溯性识别采用量化指标评估,生成任务则通过人工评估。LLMs的性能高度依赖任务,表现从中等到高不等,且没有单一模型始终优于其他模型,这表明有效应用取决于针对每个任务选择模型和提示策略。本文的贡献包括:(i)首个涵盖五项RE相关活动的LLMs跨任务实证评估;(ii)支持可复现性的复现材料;(iii)对当前LLMs用于RE的能力、局限性和实际可用性的更广泛理解。
英文摘要:
Requirements-related information is scattered across heterogeneous artefacts such as user feedback, developer discussions, and software repositories, making the extraction of actionable requirements knowledge labour-intensive and hard to scale. Large Language Models (LLMs) can support many Requirements Engineering (RE) activities, from classification and traceability identification to specification and explanation generation, but existing evidence is fragmented across tasks, artefact types, and evaluation settings, and studies rarely offer cross-task evaluations or replication packages. We present two complementary empirical studies evaluating LLMs across five RE-related activities. The first is a controlled experiment on five lightweight open-source LLMs for feedback-driven requirements classification and specification generation. The second is an exploratory industrial case study on two frontier LLMs for traceability link identification and traceability explanation generation using real project artefacts. Classification and traceability identification were assessed with quantitative metrics, and generation tasks through human evaluation. LLM performance is strongly task-dependent, ranging from moderate to high, and no single model consistently outperformed the others, indicating that effective adoption depends on selecting models and prompting strategies per task. Our contributions are: (i) the first cross-task empirical evaluation of LLMs spanning five RE-related activities, (ii) replication materials supporting reproducibility, and (iii) a broader understanding of the capabilities, limitations, and practical readiness of current LLMs for RE.