arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用智能体AI在爱立信进行情境化、多方面的代码审查

Using Agentic AI for contextualized and multifaceted code review at Ericsson

Muhammad Laiq, Ricardo Britto, Muhammad Usman, Nishrith Saini, Deepika Badampudi

arXiv 2609.15877首次发表:更新:

发表机构

Blekinge Institute of Technology; Ericsson AB(布莱金厄理工学院; 爱立信公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有代码审查方法缺乏项目情境化知识的问题,提出一种多智能体解决方案,结合专门技能与情境知识,在爱立信工业环境中实现四维度代码审查,准确率达96%,约69%识别问题被评定为重要。

AI 中文摘要

背景:由于软件系统日益复杂以及AI编码智能体加速代码生成,进行有效的代码审查变得越来越具有挑战性。基于LLM的代码审查方法在识别缺陷和提高代码质量方面显示出有前景的结果。然而,现有方法很少考虑项目特定的情境化知识,且很少有方法在工业环境中得到评估。目标:在本研究中,我们提出了一种基于多智能体的解决方案,对代码更改提供多方面的评估。方法:遵循设计科学研究流程,我们在工业环境中开发并评估了我们的解决方案。我们的解决方案将专门的智能体技能与特定情境的知识相结合,以识别代码更改在四个维度上的反模式:可读性、可维护性、可靠性和性能。使用我们的解决方案,我们为几个代码提交生成了审查意见,并识别出超过200个问题。随后,案例公司的开发人员对这些问题的正确性和重要性进行了人工验证。结果:评估结果表明,我们的解决方案在所调查的代码提交中正确识别问题的准确率达到96%。此外,在正确识别的问题中,约69%被评定为重要,其中约33%被评定为必须修复的严重问题,36%被评定为应该修复的重要问题。开发人员的定性反馈证实了这些发现,并强调了所生成审查意见的有用性。结论:我们的研究结果提供了来自工业评估的经验证据,表明将专门的智能体技能与特定情境的知识相结合,能够产生准确、实用的代码审查。

英文摘要

Context: Conducting effective code reviews is increasingly challenging due to the growing complexity of software systems and the accelerated code generation by AI coding agents. LLM-based approaches for code reviews have shown promising results in identifying defects and improving code quality. However, existing approaches rarely consider project-specific contextualized knowledge, and few have been evaluated in industrial settings. Objective: In this study, we propose a multi-agent-based solution that provides multifaceted assessments of code changes. Method: Following the Design Science Research Process, we developed and evaluated our solution in an industrial setting. Our solution combines specialized agent skills with context-specific knowledge to identify antipatterns in code changes across four dimensions: readability, maintainability, reliability, and performance. Using our solution, we generated reviews for several code commits and identified more than 200 issues. These issues were then manually validated by the developers of the case company for their correctness and importance. Results: The evaluation results show that our solution achieves 96% accuracy in correctly identifying issues in the investigated code commits. Furthermore, around 69% of the correctly identified issues were rated as important, with approximately 33% rated as severe issues that must be fixed and 36% as important issues that should be fixed. Qualitative feedback from developers corroborates these findings and highlights the usefulness of the generated reviews. Conclusion: Our findings provide empirical evidence from an industrial evaluation that combining specialized agent skills with context-specific knowledge yields accurate, practically useful code reviews.

CommentsAccepted at the 27th International Conference on Product-Focused Software Process Improvement (PROFES 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑