arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型在软件工程系统综述中的筛选性能是否已停滞?

Has LLM Screening Performance Stalled in Software Engineering Systematic Reviews?

Aleksi Huotala, Miikka Kuutila, Mika Mäntylä

arXiv 2610.10633首次发表:更新:

发表机构

University of Helsinki; LUT University(赫尔辛基大学; 拉彭兰塔-拉赫蒂理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究以SESR-Eval等为数据,评估8款新LLM的软件工程系统综述筛选性能,发现其仅略优于旧模型,尚未能替代人工,提出三类未来研究方向。

AI 中文摘要

系统综述(SR)中的筛选工作为手动完成,耗时且繁琐。已有研究探索用大型语言模型(LLM)实现该步骤的自动化,但LLM发展迅速,早期的性能结论可能已无法准确反映其筛选表现。本研究使用现有软件工程系统综述筛选基准SESR-Eval作为数据,还通过幂次采样构建了新的小型数据集SESR-Eval-Mini,可降低评估成本。基于这些数据,本研究评估了8款新LLM的筛选性能,同时测试了不同提示词、分析LLM在筛选决策与标准上的一致性,还探究了纳入与排除标准的优化对筛选性能的影响。结果显示,8款新LLM的表现仅略优于7款旧LLM:针对二次研究的平均马修斯相关系数(MCC)从0.347升至0.365;二次研究间的差异仍大于LLM间的差异;基于标准级决策计算整体筛选决策仅使性能略有下降;LLM在对应筛选决策上整体一致(平均Gwet's AC1=0.830),但部分纳入与排除标准的分歧更大;优化纳入与排除标准仅小幅提升了召回率,且仅对部分LLM简化了决策,整体影响有限。研究结论为LLM尚未能替代人工完成论文筛选,新的、成本更高的模型带来的优势十分有限,基于智能体的方法、提示词工程及进一步的标准优化是未来的三个潜在研究方向。

英文摘要

Screening in systematic reviews (SRs) is manual and time-consuming. Prior work has explored large language models (LLMs) for automating this step, but LLMs are evolving rapidly, so earlier performance claims may no longer accurately reflect their screening performance. We used an existing software engineering SR screening benchmark (SESR-Eval) as our data. We also power-sampled a new, smaller dataset (SESR-Eval-Mini) that allows evaluation at lower costs. Using this data, we evaluated eight new LLMs for screening performance. Additionally, we tested different prompts, analyzed LLM agreement in screening decisions and criteria, and examined the effect of refining the inclusion and exclusion criteria on screening performance. The eight new LLMs performed marginally better than the seven old ones: avg. MCC across secondary studies rose from 0.347 to 0.365. Differences between secondary studies are still bigger than between LLMs. Computing the overall screening decision from criterion-level decisions degraded screening performance only slightly. LLMs generally agree with each other in their corresponding screening decisions (mean Gwet's AC1 = 0.830), though certain inclusion and exclusion criteria showed larger disagreement than others. Refining the inclusion and exclusion criteria slightly improved recall and made decisions easier for some LLMs, but overall impacts of criteria refinement were modest. LLMs are not yet ready to replace humans in paper screening and the advantages new, more costly models bring, appear to be very limited. Agent-based approaches, prompt engineering, and further criteria refinement are three potential future research avenues.

Comments53 pages, four external figures available in the research artifact

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑