使用和评估生成式人工智能工具以支持系统文献综述的初步指南
Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews
- Keele University(基尔大学)
- Universidad de la República(共和国大学)
- Wrocław University of Science and Technology(弗罗茨瓦夫理工大学)
- University of Calgary(卡尔加里大学)
- Brunel University of London(伦敦布鲁内尔大学)
- Durham University(杜伦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究探讨如何用GenAI支持系统文献综述,通过快速审查、思想实验等方法,识别评估GenAI用于SLRs的问题及过程要点,形成GUEST建议,助研究人员利用GenAI开展可靠综述与评估研究,强调其需人工监督及能提供成本效益帮助。
AI中文摘要:
背景:生成式人工智能(GenAI)和大语言模型(LLMs)越来越多地用于软件工程及其他领域的学术任务,包括系统文献综述(SLRs)。然而,虽然它们能够总结文本,但不能保证能满足SLRs所需的严谨性、可靠性和透明度。目标:支持打算使用GenAI进行SLRs的研究人员,或那些进行实证研究以评估GenAI对SLR任务支持程度的人员。方法:首先,进行快速审查以识别提出评估和使用GenAI及LLMs支持SLRs指南的研究。其次,利用思想实验、文献中的相关指导以及自身进行SLRs和评估工具的经验,为在SLRs背景下如何使用和评估GenAI制定建议。结果:讨论了研究人员在评估GenAI用于SLRs时面临的问题。识别并解释了在计划、进行和报告使用GenAI的SLRs以及GenAI工具评估时要考虑的过程问题。最后,将结果总结为一组过程建议,即GUEST(GenAI在SLR任务中的使用和评估)。结论:认为GenAI需要人工监督,目前无法进行无监督的系统研究。然而,它为一些重复性任务和一些复杂任务的额外验证提供了具有成本效益的帮助。GUEST建议应有助于软件工程研究人员使用GenAI进行并报告可靠的SLRs,并提供严格的独立评估研究。
英文摘要:
Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond, including systematic literature reviews (SLRs). However, while capable of summarizing text, there is no guarantee they can meet the rigour, reliability, and transparency that SLRs require. Objectives: To support researchers intending to conduct SLRs using GenAI or those conducting empirical studies evaluating how well GenAI supports SLR tasks. Methods: First, we conducted a rapid review to identify studies that propose guidelines for evaluating and using GenAI and LLMs to support SLRs. Second, we drew on thought experiments, relevant guidance from the literature, and our own experience conducting SLRs and evaluating tools to develop recommendations for how to use and assess GenAI in the context of SLRs. Results: We discuss the problems researchers face when evaluating GenAI for SLRs. We identify and explain process issues to consider when planning, conducting, and reporting both SLRs using GenAI and evaluations of GenAI tools. Finally, we summarize our results as a set of process recommendations, which we name GUEST (GenAI Use and Evaluation in SLR Tasks). Conclusion: We argue that GenAI requires human oversight and is not currently capable of unsupervised systematic studies. However, it offers the prospect of cost-effective assistance for some repetitive tasks and for additional validation of some complex tasks. Our GUEST recommendations should help software engineering researchers both to conduct and report trustworthy SLRs using GenAI and to provide rigorous independent evaluation studies.