CheckerBench:长时程智能体能否合成静态分析检查器?
CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?
浏览论文内容
中文总结 AI 辅助
本文提出CheckerBench基准和CheckerLab评估框架,测试长时程智能体合成静态分析检查器的能力,发现当前模型平均Pass@1仅32.30%,可靠检查器开发仍具挑战。
中文摘要 AI 辅助
静态分析检查器的合成要求智能体解释缺陷规范、检查代码仓库、实现分析器特定逻辑,并通过反复的编译和分析反馈来改进检查器。现有的编码智能体基准侧重于补丁生成或漏洞检测等任务,很少评估智能体能否从零开始在代码仓库中开发出可工作的检查器。我们引入了CheckerBench,这是一个可执行的基准,包含从167个代码仓库、85个CWE和五个语言生态系统的297个CVE中衍生的300个任务。每个任务包含易受攻击和修复后的版本、固定的分析环境以及检查器脚手架。我们进一步引入了CheckerLab,这是一个通用的评估框架,独立重建提交的检查器,并测量易受攻击与修复后的诊断对比、补丁定位、误报和工具使用情况。在21种模型-工具配置和每种配置三次独立重复中,平均Pass@1为32.30%,而最佳达到45.33%。这些结果表明,对于当前的编码智能体来说,开发可靠且可复用的检查器仍然具有挑战性。
英文摘要
Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection and rarely assess whether an agent can develop a working checker in a repository from start to finish. We introduce CheckerBench, an executable benchmark of 300 tasks derived from 297 CVEs across 167 repositories, 85 CWEs, and five language ecosystems. Each task includes vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold. We further introduce CheckerLab, a common evaluation framework that independently rebuilds submitted checkers and measures vulnerable-fixed diagnostic contrast, patch localization, false positives, and tool use. Across 21 model-harness configurations and three independent repeats per configuration, mean Pass@1 is 32.30%, while the best reaches 45.33%. These results show that reliable, reusable checker development remains challenging for current coding agents.
发表机构
- East China Normal University(华东师范大学)
- Humanlaya Data(Humanlaya数据公司)
- Shanghai Jiao Tong University(上海交通大学)
- Peking University(北京大学)
- Shanghai Innovation Institute(上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。