REFINE:一种用于证据引导代码重构的多智能体大语言模型方法
REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring
浏览论文内容
中文总结 AI 辅助
本研究提出多智能体方法REFINE,结合静态分析、LLM等技术生成Java代码重构候选,在450个文件上使异味降低超68%,效果优于直接提示基线,但存在残留风险需人工审查。
中文摘要 AI 辅助
大语言模型(LLMs)为自动化代码重构提供了新机遇,但生成的变更必须减少目标质量问题,同时不引入新问题或改变与行为相关的代码结构。我们推出REFINE(Refactoring with Evidence-aware Flow for Integrated ageNtic Execution),这是一种与工具无关、具备证据感知能力的多智能体方法,用于生成Java文件级别的重构候选方案。REFINE结合了静态分析引导的异味识别、异味感知规划、基于LLM的转换、自动重分析、保留检查以及结构化报告。我们在来自15个开源系统的450个Java文件上对REFINE进行评估,使用OpenAI GPT-5.5、Google Gemini 3.1 Pro Preview和Anthropic Claude Opus 4.8生成了1350个模型通过的输出。REFINE在三种配置下分别将检测到的代码异味降低了68.26%、72.79%和68.49%,其中主要异味的降低幅度最大。匹配的150个文件直接提示基线实验显示,REFINE在编辑量更小、公共方法移除更少的情况下,实现了更高的中位数代码异味降低。不过,更广泛的质量改进并不稳定,且保留检查显示存在残留风险,包括断言/失败调用变更和公共方法移除。因此,REFINE的输出应被视为重构候选方案,在应用于仓库或系统级设置前,需要进行编译、测试、依赖分析和人工审查。
英文摘要
Large Language Models (LLMs) offer new opportunities for automated code refactoring. However, generated changes must reduce targeted quality problems without introducing new issues or altering behaviour-relevant code structures. We introduce REFINE (Refactoring with Evidence-aware Flow for Integrated ageNtic Execution), a tool-agnostic, evidence-aware multi-agent approach for generating Java file-level refactoring candidates. REFINE combines static-analysis-guided smell identification, smell-informed planning, LLM-based transformation, automated re-analysis, preservation checks, and structured reporting. We evaluate REFINE on 450 Java files from 15 open-source systems, producing 1,350 model-pass outputs using OpenAI GPT-5.5, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.8. REFINE reduces detected code smells by 68.26%, 72.79%, and 68.49% across the three configurations, respectively, with the strongest reductions observed for major smells. A matched 150-file direct-prompt baseline shows that REFINE achieves a higher median code-smell reduction with smaller edits and fewer public-method removals. However, broader quality improvements are inconsistent, and preservation checks reveal residual risks, including assert/fail-call changes and public-method removal. Therefore, REFINE outputs should be treated as refactoring candidates requiring compilation, testing, dependency analysis, and human review before adoption in repository- or system-level settings.
发表机构
- Tampere University(坦佩雷大学)
- University of Derby(德比大学)
机构由 AI 辅助整理,请以论文原文为准。