使用LLM改造代码以支持异常行为
Retrofitting Code Using LLMs to Support Exceptional Behavior
浏览论文内容
中文总结 AI 辅助
提出EXCODER框架,结合静态与动态程序分析及上下文工程,利用LLM为现有代码自动生成异常相关代码(ERC),在304个Java方法上实现85.92%的pass@1,显著超越基线。
中文摘要 AI 辅助
异常相关代码(ERC),包括抛出语句、保护这些抛出语句的条件(if语句)以及try/catch块,是软件系统的重要组成部分,它使开发人员能够检测和处理偏离预期程序行为的异常状态。然而,在大型代码库中手动编写ERC非常繁琐。我们提出了一个新任务:为现有代码改造添加ERC。即,给定代码(不含ERC)和异常行为测试(EBTs)(例如,检查如果向参数传入null值,方法是否抛出InvalidArgumentException),我们旨在自动生成缺失的ERC,使得给定的测试通过。我们设计并实现了Exception Coder(EXCODER),它通过上下文工程帮助大型语言模型(LLMs)完成此任务。EXCODER将静态和动态程序分析与LLMs相结合,向LLMs提供提取的上下文信息。为了评估EXCODER,我们构建了一个基于GitHub Java仓库的基准,系统性地从75个项目的304个方法中移除了ERC。我们的结果表明,EXCODER为自动化代码生成中的这一问题提供了一种有效但不完美的解决方案,为开发人员提供了首次以测试驱动开发方式实现ERC的途径。当与Qwen 2.5 Coder 32b结合时,EXCODER在开发者编写的测试套件上分别实现了pass@1、5和10的85.92%(比基线高12.56个百分点)、86.18%(比基线高12.82个百分点)和86.51%(比基线高13.15个百分点)的通过率。我们对生成代码的人工检查进一步揭示了EXCODER的局限性,为未来的工作指明了方向。
英文摘要
Exception Related Code (ERC), which includes throw statements, conditions (if statements) that guard those throw statements, and try/catch blocks, is an essential component of software systems, allowing developers to detect and handle exceptional states that deviate from the expected program behavior. However, manually writing ERC across large codebases is tedious. We propose a novel task: retrofitting existing code with ERC. Namely, given code (without ERC) and Exceptional Behavior Tests (EBTs) (e.g., check if method throws InvalidArgumentException if null is given as the value to the argument) we aim to automatically generate missing ERC, such that the given tests pass. We design and implement Exception Coder (EXCODER) that performs context engineering to help Large Language Models (LLMs) tackle this task. EXCODER integrates static and dynamic program analysis with LLMs by providing the extracted contextual information to the LLMs. To evaluate EXCODER, we build a benchmark constructed from GitHub Java repositories, where we systematically remove ERC in 304 methods from 75 projects. Our results demonstrate that EXCODER provides an effective, though imperfect, solution to this problem in automated code generation, offering developers the first way to implement ERC following test-driven development. When combined with Qwen 2.5 Coder 32b, EXCODER achieves pass@1, 5, and 10 rates of 85.92% (12.56 percentage points over baseline), 86.18% (12.82 p.p. over baseline), and 86.51% (13.15 p.p. over baseline), respectively, on developer-written test suites. Our manual inspection of the generated code further reveals limitations of EXCODER, pointing to directions for future work.
发表机构
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
- Cisco Systems(思科系统公司)
机构由 AI 辅助整理,请以论文原文为准。