arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估代码推荐系统:综述

Evaluating Code Recommender Systems: A Review

Daniel Borst, Stefan Sobernig

arXiv 2609.30351首次发表:更新:

发表机构

Institute for Complex Networks, WU Vienna; WU Vienna(复杂网络研究所,维也纳经济大学; 维也纳经济大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述通过系统文献回顾(2017-2024,92篇研究)评估代码推荐系统,发现评估以离线、系统为中心为主,用户中心评估稀缺,且集中于软件构建阶段,指出评估质量单一、类型不平衡的问题。

AI 中文摘要

背景:代码推荐系统(CRS)是专门化的软件系统,它们作用于源代码工件,在软件开发的所有阶段向软件开发者提供自动生成的推荐。其目标是提高软件质量,同时增强开发者的效率、有效性和体验。问题:尽管其重要性日益增长,但关于以人为中心的方式评估这些系统的现状知之甚少。研究方法:我们进行了一项系统性文献综述,以识别和综合评估CRS(2017-2024年)的主要研究。共纳入92篇出版物,并对其进行了系统性内容分析。结果:我们的研究证实,离线评估是最常见的以系统为中心的评估类型,而以用户为中心的评估(在线评估、用户研究)很少被报告。评估集中在软件构建阶段。大多数研究明确报告了有效性威胁,其中外部威胁对研究材料最为常见。结论:CRS评估的现状狭窄地聚焦于单一或少数系统质量。以系统为中心的评估占多数与以用户为中心的评估占少数之间存在不平衡。结合多种类型的评估很少被报告。

英文摘要

Context: Code recommender systems (CRSs) are specialized software systems operating on source artifacts to provide automatically generated recommendations to software developers in all phases of development. The goal is to improve software quality while enhancing developers' efficiency, effectiveness, and experience. Problem: Despite the growing importance, little is known about the state of evaluating these systems in a human-centric manner. Research Approach: We conducted a systematic literature review to identify and to synthesize primary studies evaluating CRSs (2017-2024). Ninety-two publications were included and subjected to a systematic content analysis. Results: Our study confirms that Offline Evaluations are the most common, system-centric evaluation type, whereas user-centric evaluations (Online Evaluations, User Studies) are rarely reported. Evaluations are concentrated on the Software Construction phase. Most studies explicitly report threats to validity, with External threats to study Materials being the most common. Conclusion: The state of evaluations on CRSs is narrowly focused on single or a few system qualities. There is an imbalance between a majority of system-centric evaluations and a minority of user-centric evaluations. Combined, multi-type evaluations are rarely reported.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑