arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估可解释人工智能对人工智能辅助代码审查中信任的影响

Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review

Zhenhan Gao, Marvin Muñoz Barón, Umm-e Habiba, Daniel Graziotin, Stefan Wagner

arXiv 2607.24601首次发表:更新:

发表机构

Technical University of Munich; University of Hohenheim(慕尼黑工业大学; 霍恩海姆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究可解释人工智能(XAI)对开发人员在人工智能辅助代码审查中信任的影响,通过对34名参与者的受试者内用户研究,比较不同XAI支持水平的系统,发现解释水平显著影响信任和认同,结果为可信代码审查系统设计等研究提供信息。

AI 中文摘要

背景:大语言模型(LLMs)越来越多地用于自动化代码审查,但其决策背后的推理仍难以理解。开发人员难以评估LLM生成的审查的有效性,难以确定对其的信任程度。可解释人工智能(XAI)在代码审查中的作用及其对信任的影响仍未得到充分探索。目的:研究XAI对开发人员在人工智能辅助代码审查中的信任的影响。方法:对34名参与者进行了一项受试者内用户研究,比较了三个具有不同XAI支持水平的基于LLM的代码审查系统:条件A(详细解释和审查反馈)、条件B(仅审查反馈)和条件C(无解释)。参与者与人工智能生成的审查一起审查实际的代码更改请求。我们测量了信任感知、对人工智能建议的认同、每个决策给出的理由以及所花费的时间。结果:解释水平显著影响对人工智能建议的信任和认同,但方式不同。完整解释(A)产生最高的感知信任(M = 3.99/5)但不是最高的认同,而适度解释(B)达到最高的认同(89.22%)。这可能表明更多的解释促使开发人员更频繁地质疑人工智能建议。无解释(C)导致最低的信任和认同。解释水平对审查时间没有显著影响。决策最常引用的理由是代码可读性和正确性。结论:将XAI纳入代码审查会显著改变信任感知和对人工智能建议的认同。这些结果为基于人工智能的可信代码审查系统的设计和评估以及人工智能辅助软件开发的人为因素研究提供了信息。

英文摘要

Background: Large language models (LLMs) are increasingly used to automate code review, but the reasoning behind their decisions remains hard to understand. Developers struggle to assess the validity of LLM-generated reviews, making it difficult to gauge how much trust to place in them. The role of Explainable AI (XAI) in code review and its impact on trust remain underexplored. Objective: We study the influence of XAI on developer trust in AI-assisted code reviews. Method: We conducted a within-subjects user study with 34 participants, comparing three LLM-based code review systems with varying levels of XAI support: Condition A (detailed explanation and review feedback), Condition B (review feedback only), and Condition C (no explanations). Participants reviewed real-world code change requests alongside the AI-generated reviews. We measured trust perceptions, agreement with the AI recommendation, the reasoning given for each decision, and the time taken. Results: The level of explanation significantly influences both trust and agreement with AI recommendations, but in different ways. Full explanations (A) yield the highest perceived trust (M = 3.99/5) but not the highest agreement, whereas moderate explanations (B) achieve the highest agreement (89.22%). This could suggest that more explanation prompts developers to question AI recommendations more frequently. No explanations (C) results in the lowest trust and agreement. Explanation level did not significantly affect review time. The most commonly cited reasons for decisions were code readability and correctness. Conclusion: Incorporating XAI into code review significantly changes trust perceptions and agreement with AI recommendations. These results inform the design and evaluation of trustworthy AI-based code review systems, as well as studies on the human factors of AI-assisted software development.

Comments23 pages, 4 figures, 5 tables. To appear in Proceedings of the ACM on Software Engineering (PACMSE), Vol. 3, No. ISSTA, Article ISSTA093 (ISSTA 2026). Published under CC BY 4.0. Replication package: https://doi.org/10.5281/zenodo.21457282

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑