arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型能否无需代码修复?面向无代码漏洞修复的自动化验证

Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes

Utku Boran Torun, Veli Karakaya, Eray Tüzün

arXiv 2610.11963首次发表:更新:

发表机构

Bilkent University(比尔肯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于执行的自动化流程,评估LLM在真实浏览器环境生成无代码修复的能力,发现其修复效果高度依赖执行器,多数修复需验证后交付用户。

AI 中文摘要

无代码修复通过指导用户更改设置、更新至已修复问题的版本或调整工作流程来解决无效漏洞报告。手动验证所提出的无代码修复是否能解决报告的漏洞需要开发者投入大量时间。本研究提出一种基于执行的自动化流程,用于评估大语言模型(LLM)在真实浏览器环境中生成无代码修复的能力。我们评估了之前研究基准发布的12种配置生成的322个无代码修复,涵盖归类为配置错误、版本错误或外部系统与依赖项的漏洞报告。执行代理会按照每个修复的自然语言说明应用该修复,特定问题检查器则确定报告的漏洞是否仍然存在。我们使用三种执行器重复该流程:两个计算机使用代理OpenCUA-72B和Claude Sonnet 5,以及一个多模态代理大语言模型Meta的Muse Glimmer。仅17.6%的候选问题能够被设置并通过两个健全性检查。在322个修复中,根据执行器的不同,有14.6%至49.7%的修复解决了漏洞,最强配置Vanilla流程中的Claude Opus 4.6在Claude Sonnet 5下的修复解决率高达74.1%。仅更换执行器就使某一配置的解决率平均变化了38.8%,且三个执行器仅在46.9%的修复上达成一致判断。与人工执行相比,这些执行器对抽样修复的人工共识匹配度为66.1%至88.1%。即使在最佳执行器下,LLM生成的无代码修复中也仅有不到一半能解决报告的漏洞,因此此类修复在交付给用户前需要验证。基于执行的验证可提供此功能,但所测能力高度依赖执行器,评估必须报告并控制这一点。

英文摘要

A no-code fix resolves an invalid bug report by directing the user to change a setting, update to a version where the problem is already fixed, or adjust their workflow. Manually verifying whether a proposed no-code fix resolves the reported bug takes considerable developer time. This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment. We evaluate 322 no-code fixes generated by the 12 configurations released with the benchmark of a previous study, covering bug reports categorized as Faulty Configuration, Wrong Version, or External System & Dependency. An executor agent applies each fix by following its natural-language instructions, and an issue-specific checker determines whether the reported bug persists. We repeat the pipeline with three executors: two Computer-Use Agents, OpenCUA-72B and Claude Sonnet 5, and one multimodal agentic LLM, Meta's Muse Glimmer. Only 17.6% of the candidate issues could be set up and passed both sanity gates. Across the 322 fixes, 14.6% to 49.7% resolved the bug depending on the executor, and the strongest configuration, Claude Opus 4.6 in the Vanilla pipeline, resolved up to 74.1% of its fixes under Claude Sonnet 5. Changing only the executor shifted a configuration's resolution rate by 38.8% on average, and the three executors reached the same verdict on only 46.9% of the fixes. Compared with human execution, the executors matched the human consensus for 66.1% to 88.1% of the sampled fixes. Even under the best executor, fewer than half of the LLM-generated no-code fixes resolve the reported bug, so such fixes need verification before they reach users. Execution-based verification can provide this, but the measured capability depends strongly on the executor, which evaluations must report and control.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑