AI 中文总结
研究程序员评估大型语言模型生成断言的能力,通过实验发现程序员判断正确断言较准,判断错误断言较差,自然语言解释未带来整体益处,凸显帮助开发者评估机器生成可靠性工件的重要性。
AI 中文摘要
代码理解和代码审查是软件工程的关键任务,随着人工智能代码生成工具的使用增加,其重要性愈发凸显。生成式人工智能有望支持这些活动,但效果未知。我们对86名Python程序员进行了对照实验及后续出声思考研究,以考察开发者评估不同质量生成断言的正确性和完整性的能力以及自然语言解释的影响。结果显示程序员判断正确断言较准确(74%准确率),但判断错误断言时表现不佳(49%准确率),且自然语言解释未带来整体益处,低质量解释还会损害评估准确性并增加开发者信心。研究表明人工智能辅助可能无法提高代码理解和审查的可靠性,强调了帮助开发者评估机器生成的可靠性工件的重要性。
英文摘要
Code comprehension and code review are already critically important software engineering tasks, and the rising use of AI code generation tools is only increasing that importance. Generative AI has the possibility of supporting these activities, for example by augmenting code with assertions and natural-language explanations describing code behavior. However, little is known about how effective such support may be. We conduct a controlled experiment with 86 Python programmers and a follow-up think-aloud study to examine developers' ability to assess the correctness and completeness of generated assertions of varying quality, and to investigate how natural-language explanations influence these assessments. While programmers can somewhat accurately judge correct assertions (74% accuracy), they perform poorly when shown incorrect assertions (49% accuracy), despite reporting similar levels of confidence in both judgments. This difference in judgment accuracy is statistically significant (p < 0.001): the odds of a developer accurately judging a correct assertion was nearly three times higher than the odds of accurately judging an incorrect assertion (OR = 2.94). Surprisingly, natural-language explanations of assertions provided no overall benefit. Furthermore, low-quality explanations could impair specification assessment accuracy (p = 0.037, OR = 0.58) while simultaneously increasing developer confidence (p = 0.005, 3.99/5 vs. 4.25/5). Our findings suggest that, contrary to common assumptions, AI assistance may not improve the reliability of code comprehension and review. More broadly, our findings highlight the importance of helping developers evaluate machine-generated reliability artifacts, in addition to generating them.
Comments10 pages, 6 figures