arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SecProbe:编码智能体在网络安全漏洞上的自适应评估

SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities

Xiaonan Luo, Yue Huang, Kehan Guo, Ping He, Chuan Zou, Chujie Gao, Lichi Li, Yuchen Ma, Zhangchen Xu, Zichen Chen, Yufei Han, Xiangliang Zhang

arXiv 2609.33763首次发表:更新:

发表机构

University of Notre Dame; Bake AI; Vanderbilt University; University of Pennsylvania; LMU Munich; University of Washington; Stanford University; Inria(圣母大学; Bake AI; 范德堡大学; 宾夕法尼亚大学; 慕尼黑大学; 华盛顿大学; 斯坦福大学; 法国国家信息与自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SecProbe结合项目反应理论与按需合成漏洞修复任务,实现自适应评估,在保持能力估计准确的同时减少任务量,高效揭示编码智能体的漏洞识别与修复差距。

AI 中文摘要

评估编码智能体的网络安全漏洞意识,需要能够揭示能力差距且随着模型演进仍保持信息量的评估方法。静态基准提供固定的覆盖范围和难度,而稀缺的漏洞仓库和昂贵的专家编写限制了其大规模更新。我们提出SecProbe,一个自适应评估框架,结合项目反应理论(IRT)与按需合成仓库规模的漏洞修复任务。根据观察到的性能,SecProbe估计智能体能力并识别额外证据最具信息量的位置,据此选择现有任务或合成新任务。作为一个用例,我们构建了涵盖六种编程语言和151种CWE类型的353个任务,并使用两个智能体框架评估九个前沿模型。成功率峰值仅为28.33%,凸显了漏洞识别和修复方面的巨大差距。与随机和一次性基线相比,SecProbe在实现相当的智能体能力估计的同时,要求智能体解决的任务数量减少多达29.5%。这些结果支持自适应评估作为评估网络安全漏洞意识的一种高效且具有区分力的方法。

英文摘要

Assessing cybersecurity vulnerability awareness in coding agents requires evaluations that reveal capability gaps and remain informative as models evolve. Static benchmarks offer fixed coverage and difficulty, while scarce vulnerable repositories and costly expert authoring limit their renewal at scale. We introduce SecProbe, a framework for adaptive evaluation that combines Item Response Theory (IRT) with on-demand synthesis of repository-scale vulnerability-repair tasks. From observed performance, \textsc{SecProbe} estimates agent ability and identifies where additional evidence is most informative, selecting existing tasks or synthesizing new ones accordingly. As one use case, we construct 353 tasks spanning six programming languages and 151 CWE types and evaluate nine frontier models with two agent harnesses. Success rates peak at 28.33\%, highlighting substantial gaps in vulnerability recognition and repair. Compared with random and one-shot baselines, \textsc{SecProbe} achieves comparable agent ability estimates while requiring agents to solve up to 29.5\% fewer tasks. These results support adaptive evaluation as an efficient and discriminative approach to assessing cybersecurity vulnerability awareness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑