arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越Pass@k:衡量智能体代码生成的可靠性与安全性

Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation

Jiajun Jiang, Sharon Zheng, Natan Vidra, Spurthi Setty

arXiv 2608.14711首次发表:更新:

发表机构

Stevens Institute of Technology; Cornell University(史蒂文斯理工学院; 康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究指出AI编码智能体基准误用Pass@k指标,提出正确的Reliability@k及安全调整版,通过实验验证误用指标的偏差,还在SWE-bench试点中发现真实场景下的核心问题。

AI 中文摘要

AI编码智能体基准使用Chen等人(2021)提出的Pass@k评估器对智能体进行排名,但当前实现存在误用:将n设置为单次提交中的单元测试数量,而非独立的rollout尝试次数,混淆了测试套件规模与尝试独立性。我们诊断了这一操作化错误,通过反例证明其存在,并提出了正确应用的Reliability@k,其中n为独立rollout次数,c为(任务,智能体)对中完全通过的rollout次数。在合成多rollout基准中,误用的指标使报告分数绝对数值升高了0.85-0.97(报告值为0.96-0.98,修正后为0.00-0.12),而廉价的单次rollout代理无法替代重复运行(斯皮尔曼相关系数ρ=0.417)。鉴于功能正确性不代表安全的证据,我们还提出了安全调整后的Reliability@k,仅统计功能正确且无高严重性不安全模式的rollout。在对三个智能体的初始实时API测试中,根据当前扫描器和阈值,该调整未改变任何排名,因此将其作为建议的补充视角,其决定性评估需要未来更大规模的运行。最后,在5个任务的SWE-bench Verified初步试点中,我们在真实仓库环境中观察到相同核心问题:宏平均隐藏测试通过率为0.80,而严格任务解决率仅为0.20。

英文摘要

AI coding agent benchmarks rank agents with the Chen et al. (2021) pass@k estimator, but current implementations misapply it: they set n to the number of unit tests in a single submission rather than the number of independent rollout attempts, conflating test-suite size with attempt independence. We diagnose this operationalization error, prove it by counterexample, and propose reliability@k, the same estimator applied correctly, with n = independent rollouts and c = fully-passing rollouts per (task, agent) pair. In a synthetic multi-rollout benchmark, the misapplied metric inflates reported scores by 0.85-0.97 in absolute terms (0.96-0.98 reported vs. 0.00-0.12 corrected), and a cheap single-rollout proxy fails to substitute for repeated runs (Spearman $ρ= 0.417$). Motivated by evidence that functional correctness does not imply security safety, we additionally propose security-adjusted reliability@k, which counts only rollouts that are both functionally correct and free of high-severity insecure patterns. In an initial live-API test with three agents, the adjustment did not change any ranking under our current scanner and threshold, so we present it as a proposed complementary lens whose decisive evaluation requires better-powered future runs. Finally, a preliminary 5-task SWE-bench Verified pilot observes the same core concern in a real repository setting: macro-averaged hidden-test pass rate was 0.80 while strict task resolution was 0.20.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑