arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18389cs.AIcs.SE

锯齿状前沿:评估代码智能体对语义保留转换的鲁棒性

A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations

Hasan Najib Mahmud, Shreya Gupta, Isha Chaudhary, Nathaniel Enis, Ravi Mangal, Gagandeep Singh, Corina Pasareanu

AI总结:

该研究评估AI代码智能体对语义保留代码转换的鲁棒性,发现其存在锯齿状鲁棒性前沿,mini-SWE agent更鲁棒,顶级模型仍易受此类扰动影响,引发部署可靠性担忧。

AI中文摘要:

AI代码智能体越来越多地被用于解决实际软件问题,然而它们在代码表层变化下的可靠性仍未得到充分理解。我们评估修复仓库级问题的代码智能体,在周围代码库被重写为语义等价形式时是否仍保持可靠。我们引入一种随机变体采样器,该采样器应用常见的语义保留转换(SPT),涵盖控制流重写、死代码注入和标识符重命名,以生成扰动变体。我们评估两种智能体框架(mini-SWE agent和OpenCode),每种框架分别由四种前沿模型(Claude Opus 4.5、Kimi K2.5、MiniMax M2.5和Qwen 3.6-27B)提供支持,测试实例来自SWE-bench Verified和SWE-bench Pro。对于每个实例,智能体在未扰动和扰动变体上运行多次,产生成对的解决率估计值,以将扰动效应与内在随机性隔离开来。我们发现,大多数配置中解决率下降幅度较小:受影响最严重的配置中平均解决率下降高达6.7个百分点,在16种模型、框架和数据集组合配置中,有6种存在统计上显著的下降。关键的是,不存在跨框架统一的鲁棒性模型排名——在SWE-bench Verified上,Qwen在mini-SWE agent下是最鲁棒的模型之一,但在OpenCode下却是最脆弱的,这揭示了锯齿状的鲁棒性前沿。更简单的框架(mini-SWE agent)对扰动更具鲁棒性。我们的结果表明,即使是顶级前沿模型也易受语义保留扰动的影响,尽管影响并不均匀,这引发了对AI代码智能体在多样化实际代码库中部署可靠性的担忧。

英文摘要:

AI code agents are increasingly deployed to resolve real software issues, yet their reliability under superficial code variations remains poorly understood. We evaluate whether coding agents that repair repository-level issues remain reliable when the surrounding codebase is rewritten into a semantically equivalent form. We introduce a random variant sampler that applies common semantics-preserving transformations (SPTs) - spanning control-flow rewrites, dead-code injection, and identifier renaming - to produce perturbed variants. We evaluate two agentic scaffolds (mini-SWE agent and OpenCode) each backed by one of four frontier models (Claude Opus 4.5, Kimi K2.5, MiniMax M2.5, and Qwen 3.6-27B) across instances drawn from SWE-bench Verified and SWE-bench Pro. For each instance, the agent is run multiple times on the unperturbed and perturbed variants, yielding paired resolve-rate estimates that isolate the perturbation effect from intrinsic stochasticity. We find small degradation in most configurations: up to 6.7 percentage points mean resolve-rate drop in the most affected configurations with statistically significant degradations in 6 of 16 configurations of model, scaffold, and dataset. Crucially, no single model ranking by robustness holds across scaffolds - Qwen is among the most robust under mini-SWE agent on SWE-bench Verified yet the most brittle under OpenCode - revealing a jagged robustness frontier. The simpler scaffold (mini-SWE agent) is more robust to perturbation. Our results demonstrate that even top frontier models are susceptible to semantics-preserving perturbations although the effect is not uniform, raising concerns about the deployment reliability of AI code agents in diverse real-world codebases.

补充信息

↑