arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00924cs.SEcs.AI

RefactorAssist:用于可靠代码重构的智能体式优化工具

RefactorAssist: Agentic Refinement for Reliable Code Refactoring

Jonathan Cordeiro, Shayan Noei, Ying Zou

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM生成代码重构的错误问题,开发了RefactorAssist智能体,通过静态修复结合测试引导的智能体修复,将重构累计通过率提升至94.2%,提高了LLM重构代码的可靠性。

中文摘要 AI 辅助

代码重构旨在改进源代码的内部结构,同时不影响其功能行为。大型语言模型(LLM)的最新进展已展现出自动化软件工程任务(如代码重构)的潜力,但LLM生成的重构代码常引入细微错误,导致功能行为变更和单元测试失败,限制了其实际应用。为解决LLM生成重构代码的局限性,我们分析其失败的根本原因并开发了RefactorAssist智能体,以提升LLM生成重构代码的功能正确性。为此,我们使用10个带有原生测试套件的开源Java项目,手动评估LLM生成的重构代码未通过单元测试的原因。我们设计了一种智能体式方法,利用单元测试日志、错误解释、项目上下文检索和代码差异来指导迭代重构。研究发现,失败的主要原因包括上下文误解/幻觉(24.3%)、不正确或不一致的重命名(15.3%)、添加新功能或变量(13.7%)、代码不完整(11.3%)、语法和结构错误(9.7%)、未处理的边界情况(9%)、类型处理不当(8.7%)以及超出作用域的变量(8%)。为使方法具有成本效益,RefactorAssist首先应用静态修复步骤处理缺失的导入、不平衡的括号和编译错误,无需使用LLM;对于剩余的失败,RefactorAssist结合错误日志和代码差异,在最佳配置下,对剩余失败的修复率可达70.8%,累计通过率为94.2%。这些结果表明,静态检查和测试引导的、感知上下文的智能体修复可提升LLM生成重构代码的可靠性,使其更接近开发者工作流中的实际集成。

英文摘要

Code refactoring aims to enhance the internal structure of source code without affecting its functional behavior. The recent advancements of Large Language Models (LLMs) have demonstrated potential for automating software engineering tasks, such as code refactoring. However, the refactorings produced by LLMs often introduce subtle errors, leading to functional behavior changes and failed unit tests, which limit their practical adoption. To address the limitations of LLM-generated refactorings, we analyze the root causes of their failures and develop the RefactorAssist agent to improve the functional correctness of LLM-generated refactorings. To this end, we use 10 open-source Java projects with their native test suites and manually evaluate why LLM-generated refactorings fail unit tests. We then design an agentic approach that leverages unit-test logs, error explanations, project context retrieval, and code diffs to guide the iterative refactoring. Our findings show that the main reasons for failure are context misunderstanding/hallucination (24.3%), incorrect or inconsistent renaming (15.3%), adding new functionality or variables (13.7%), code incompleteness (11.3%), syntax and structural errors (9.7%), edge cases not handled (9%), improper type handling (8.7%), and variables outside scope (8%). To make our approach cost-effective, RefactorAssist first applies a static repair step for missing imports, unbalanced brackets, and compilation errors without LLMs. For remaining failures, RefactorAssist incorporates error logs and code diffs, achieving up to a 70.8% repair rate on the remaining failures and a 94.2% cumulative pass rate under the best-performing configuration. These results indicate that static checks and test-guided, context-aware agentic repair can increase the reliability of LLM-generated refactorings, bringing them closer to practical integration within developer workflows.

发表机构

  • Queen's University(女王大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑