AI 中文总结
BreakGuard是一种基于大语言模型生成测试的方法,可静态提取客户端焦点方法生成测试,在89个真实破坏性变更上检测到30.3%的变更,对崩溃型破坏性变更检测可靠性更高。
AI 中文摘要
开源库在软件开发中发挥着重要作用,提供可复用功能以加快开发进程。随着库的演进,它们会发布新增功能、修复漏洞或应用安全补丁的新版本,在此过程中可能引入破坏性变更(BCs),改变运行时行为并破坏客户端应用,从而打破与客户端建立的约定。客户端测试套件通常无法检测到这些破坏性变更,原因在于库的覆盖范围有限,无法覆盖客户端代码库中使用的所有库方法。我们提出了BreakGuard,一种用于检测客户端破坏性变更的方法:BreakGuard会静态提取每个调用目标库方法(调用点)的客户端方法(焦点方法),然后为每个焦点方法生成测试;若测试在破坏性前版本通过但在破坏性版本失败,则该测试可检测到破坏性变更。我们在BUMP数据集中的89个真实破坏性变更上评估了该方法,使用3种大语言模型(GPT-4o、Qwen3-coder-480B、GPT-OSS-120B)和3种上下文级别(最小、方法、类)。采用性能最佳的配置时,BreakGuard可检测到30.3%的破坏性变更(89个中的27个),每个检测到的破坏性变更的平均成本约为0.90美元。我们成功检测到不同库类别(如JSON库、日志、解析)的破坏性变更,但发现大语言模型生成的测试对检测崩溃型破坏性变更的可靠性高于行为型破坏性变更。
英文摘要
Open-source libraries play an important role in software development by providing reusable features that expedite the development process. As libraries evolve, they release new versions that add features, fix bugs, or apply security patches. In this process, they may break the contract established with their clients by introducing breaking changes (BCs) that alter the runtime behavior and break client applications. Client-side test suites often fail to detect these BCs because of limited library coverage that does not exercise all library methods used in the client's codebase. We propose BreakGuard, an approach that generates a test suite to detect breaking changes in clients. BreakGuard statically extracts every client method (focal method) that invokes the target library method (call site), then generates tests per focal method. A test detects a BC if it passes on the pre-breaking version and fails on the breaking version. We evaluate our approach on 89 real-world breaking changes from the BUMP dataset, using 3 LLMs (GPT4o, Qwen3-coder-480B, GPT-OSS-120B) and three context levels: minimal, method, and class. Using the best-performing configuration, BreakGuard detects 30.3% of breaking changes (27 of 89) at a mean cost of roughly $0.90 USD per detected breaking change. We successfully detected BCs from different library categories (e.g., JSON libraries, logging, parsing), but we find LLM-generated tests to be more reliable for detecting crash-type breaking changes as opposed to behavioural BCs.