arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新审视反馈驱动的大语言模型(LLM)代码修复:一项复现及探索性Java扩展研究

Revisiting Feedback-Driven LLM Code Repair: A Replication and Exploratory Java Extension

Louis Lalonde, Wassim Keddache, Thomas Perron Touchette, Leuson Da Silva, Foutse Khomh

arXiv 2609.00362首次发表:更新:

AI 中文总结

该研究复现了FeedbackEval基准的Python代码修复实验,扩展构建了Java代码修复基准,发现Python与Java场景下反馈有效性排名存在差异,提示需开展更严谨的多语言评估。

AI 中文摘要

自大型语言模型(LLM)问世以来,从业者越来越多地利用它们来支持软件工程任务,包括自动代码修复,取得了良好的效果。然而,关于可复现性和可泛化性的问题在很大程度上仍未得到探索。为了进一步评估这些问题及其相关影响,我们对FeedbackEval基准测试[1]进行了部分复现并开展了探索性Java扩展,该基准测试用于评估LLM如何利用不同反馈类型进行Python代码修复。首先,我们使用GPT-4o和Claude 3.5 Sonnet在394个修复任务上部分复现了原始研究,复现并观察到了原始研究报告的主要定性趋势。其次,我们通过从50个Java任务中构建100个错误修复实例并评估反馈有效性,开展了探索性Java扩展。我们的结果表明,来自Python的先前结论可能对基准构建、反馈表示和工具生态系统敏感,这促使人们开展更具控制性的多语言基准测试。具体而言,尽管在我们的Python复现中测试反馈仍是最强的反馈类型,但在我们的Java扩展中并未观察到相同的排名,因为简单的和基于JUnit的测试反馈没有显著差异。我们假设反馈信息量和工具生态系统的差异,例如测试框架的冗长程度,可能是造成这种差异的部分原因。最后,更简洁的提示可降低成本,且修复效果无显著差异。总体而言,我们的发现在部分受控复现下确认了关键趋势,并强调在基于LLM的修复系统中需要更严格的多语言评估和精心的反馈设计。

英文摘要

Since the advent of Large Language Models (LLMs), practitioners have increasingly leveraged them to support their software engineering tasks, including automated code repair, showing promising results. Yet, concerns regarding reproducibility and generalizability remain largely unexplored. To further evaluate these concerns and associated impacts, we partially reproduce and conduct an exploratory Java extension of the FeedbackEval benchmark [1], which evaluates how LLMs leverage different feedback types for Python code repair. First, we partially replicate the original study on 394 repair tasks using GPT-4o and Claude 3.5 Sonnet, reproducing and observing the main qualitative trends reported in the original work. Second, we conduct an exploratory Java extension by constructing 100 erroneous repair instances from 50 Java tasks and evaluating feedback effectiveness. Our results show that previous conclusions from Python may be sensitive to benchmark construction, feedback representation, and tooling ecosystem, motivating more controlled multilingual benchmarks. Specifically, while test feedback remains the strongest feedback type in our Python replication, the same ranking is not observed in our Java extension, as simple and JUnit-based test feedback do not differ significantly. We hypothesize that differences in feedback informativeness and tooling ecosystems, such as the verbosity of test frameworks, may partly explain such a difference. Finally, lighter prompts reduce cost without significant differences in repair effectiveness. Overall, our findings confirm key trends under a partially controlled replication and highlight the need for more rigorous multilingual evaluation and careful feedback design in LLMbased repair systems.

Comments10 pages, 2 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑