arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21298cs.SE

需求工程中的人机协作:大语言模型(LLM)对需求检查的负面影响证据

Human-AI Collaboration in Requirements Engineering: Evidence of the Negative Effect of LLMs on Requirements Inspection

Giovanna Broccia, Julian Frattini, Chetan Arora, Maurice H. ter Beek, Alessandro Fantechi, Andreas Vogelsang, Alessio Ferrari

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过控制交叉实验发现,大语言模型(LLM)支持会降低新手需求检查员的气味检测准确率,且先使用LLM学习需求检查会减缓技能获取。

中文摘要 AI 辅助

背景:需求检查(RI)是软件生命周期早期检测需求制品潜在缺陷的成熟实践,近期大语言模型(LLM)的进展激发了其支持需求工程(RE)任务的兴趣,但关于LLM作为人类RI协作助手的效果的实证证据仍匮乏。目标:研究LLM支持对人类执行的RI的影响,从气味识别、严重程度分类(有害vs无害)的检查有效性及检查时长两方面考量。方法:开展包含34名参与者的控制交叉设计实验,参与者在有、无LLM支持的情况下检查文本规格,识别并分类需求气味,同时记录检查时间,对每个结果变量采用贝叶斯回归模型分析,考虑交叉设计引发的有效性威胁及协变量、中介变量。结果:LLM支持对气味检测准确率有负面影响,但对气味分类或任务时长无显著影响;实验周期间存在学习效应,但先使用LLM支持执行RI时,该效应会减弱。结论:研究提供实证证据,表明LLM支持不一定提升性能,反而可能阻碍新手检查员的表现,且从一开始就使用LLM支持学习RI可能减缓技能获取过程,对LLM支持下的学习构成威胁。

英文摘要

Background. Requirements inspection (RI) is a well-established practice for detecting potential defects in requirements artifacts early in the software lifecycle. Recent advances in large language models (LLMs) have stimulated interest in their potential to support requirements engineering (RE) tasks. However, empirical evidence on the effects of LLMs when used as collaborative assistants in human-performed RI remains scarce. Aims. We aim to investigate the impact of LLM support on human-performed RI, considering inspection effectiveness in terms of smell identification and severity classification (i.e., nocuous vs innocuous), as well as inspection duration. Method. We conducted a controlled crossover design experiment with 34 participants, who inspected textual specifications with and without LLM support, identifying and classifying requirements smells while recording inspection time. We analyzed the data using one Bayesian regression model per outcome variable, accounting for validity threats induced by the crossover design as well as covariates and mediators. Results. Results show that LLM support negatively affects smell detection accuracy but has no significant effect on smell classification or task duration. A learning effect is present across experimental periods, but reduced when RI is first performed with LLM support. Conclusions. Our findings provide empirical evidence that LLM support does not necessarily improve performance and may, instead, hinder it for novice inspectors. Moreover, the results suggest that learning RI with LLM-support from the beginning may slow down the skill acquisition process, implying threats for LLM-supported learning.

↑