AI 中文总结
本文以Simulink/Stateflow建模的CPS为研究对象,扩展FlowRepair方法用LLM生成的变异替换部分算子,受控实验发现该方法大幅降低修复性能,分析指出盲目集成LLM到基于搜索的APR存在根本局限,需发展混合方法。
AI 中文摘要
基于搜索的自动化程序修复(APR)技术依赖精心设计的变异算子来探索候选修复方案的空间。大型语言模型(LLMs)的最新进展表明,生成式模型可以通过动态提出修复方案来替代这类算子。在本文中,我们在以Simulink/Stateflow建模的网络物理系统(CPS)背景下研究这一假设。我们扩展了最先进的FlowRepair方法,用LLM生成的变异替换其部分变异算子,以实现更灵活、更具表达力的补丁生成。我们在包含四个CPS领域的19个真实故障Stateflow模型的基准上评估该方法,使用与FlowRepair相同的实验设置,在相同的时间预算下进行受控比较。与预期相反,在该受控评估中,基于LLM的变异在FlowRepair实验设置下大幅降低了修复性能。在测试的LLM变体中,基于LLM的修复为4至6个模型生成了合理补丁,为4个模型生成了有效补丁,而原始方法分别为18个和16个。我们的分析显示,在这种集成中,LLM难以进行精确的符号编辑,缺乏行为反馈,并产生了阻碍有效探索的嘈杂搜索空间。这些发现并非表明LLM在APR中的普遍局限性,而是突出了将LLM盲目集成到基于搜索的APR中的根本性限制,并推动了将结构化变异与生成式引导相结合的混合方法的发展。
英文摘要
Search-based Automated Program Repair (APR) techniques rely on carefully designed mutation operators to explore the space of candidate fixes. Recent advances in Large Language Models (LLMs) suggest that generative models could replace such operators by dynamically proposing repairs. In this paper, we investigate this hypothesis in the context of Cyber-Physical Systems (CPSs) modeled in Simulink/Stateflow. We extend the state-of-the-art FlowRepair approach by replacing a subset of its mutation operators with LLM-generated mutations, enabling more flexible and expressive patch generation. We evaluate the approach on a benchmark of 19 real-world faulty Stateflow models across four CPS domains, using the same experimental setup as FlowRepair for controlled comparison under the same wall-clock budget. Contrary to expectations, in this controlled evaluation, the LLM-based mutation substantially degrades repair performance under the FlowRepair experimental setup. Across the tested LLM variants, the LLM-based repair produced plausible patches for 4-6 models and valid patches for 4 models, compared to 18 and 16, respectively, with the original approach. Our analysis reveals that, in this integration, LLMs struggle with precise symbolic edits, lack behavioral feedback, and generate a noisy search space that hinders effective exploration. Rather than showing a general limitation of LLMs for APR, these findings highlight fundamental limitations of naively integrating LLMs into search-based APR and motivate hybrid approaches that combine structured mutation with generative guidance.