架构之前的基线:评估用于自主渗透测试的编码代理
Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing
浏览论文内容
中文总结 AI 辅助
研究针对自主渗透测试,以默认编码CLI代理为基线,在XBOW基准上用不同模型运行,比较不同框架结果及模型扩展效果,发现专业框架有提升,普通代理也能解决不少问题,新模型可提升同一框架,强调未来评估应报告模型匹配的普通代理基线。
中文摘要 AI 辅助
近期自主渗透测试论文在前沿语言模型周围添加多组件安全框架时报告了高基准分数。由于这些系统常改变架构和基础模型,难以判断性能源于框架还是基础模型。本文对104任务的XBOW基准进行了对照研究,使用默认编码CLI代理作为普通代理基线。先以相同GPT-5模型、预算等运行Codex、OpenCode和Pi,确定最强同模型基线并测试特定安全提示变体能否提高分数。接着在最接近的可用模型匹配下比较默认Codex框架与已发表的MAPTA和PentestGPT V2结果。最后用GPT-5.2和GPT-5.5重复普通代理实验以衡量同一框架内模型扩展。结果显示情况复杂但实用。专业框架能提高基准分数并改善成本效率,但普通编码代理已能解决大部分基准问题;重复普通代理运行在联合覆盖方面可匹配或超过一些已发表的架构分数,新模型能显著提升同一框架。未来评估在将基准提升仅归因于架构设计前应报告模型匹配的普通代理基线。
英文摘要
Recent autonomous penetration testing papers report high benchmark scores while adding multi-component security harnesses around frontier LLMs. Because these systems often change both architecture and backbone model, it is difficult to tell how much performance comes from the harness rather than from the underlying model. This paper presents a controlled study on the 104-task XBOW benchmark using default coding CLI agents as plain-agent baselines. We first run Codex, OpenCode, and Pi with the same GPT-5 model, budget, target interface, and scoring rule. This phase identifies the strongest same-model baseline and tests whether security-specific prompt variants improve its observed score. We then compare the default Codex scaffold with published MAPTA and PentestGPT V2 results under the closest available model matches. Finally, we repeat the plain-agent experiment with GPT-5.2 and GPT-5.5 to measure model scaling inside the same scaffold. The results show a mixed but practical picture. Specialised harnesses can add measurable benchmark lift and may improve cost efficiency, but plain coding agents already solve a large share of the benchmark; repeated plain-agent runs can match or exceed some published architecture scores in union coverage, and newer models substantially improve the same scaffold. Future evaluations should report model-matched plain-agent baselines before attributing benchmark gains to architecture design alone.
发表机构
- Kroda Labs(Kroda实验室)
机构由 AI 辅助整理,请以论文原文为准。