AI 中文总结
研究通过成本-成功率视角评估语言模型安全代理,在攻防挑战中比较固定成本下模型性能,发现红蓝队任务扩展模式不同,进攻性CTF性能与测试时计算量有关,防御性SOC调查更依赖工具使用等,强调安全代理基准应考量经济效率与运营契合度。
AI 中文摘要
安全代理评估通常在充足的推理预算下衡量峰值攻击能力,侧重于漏洞发现、利用开发、渗透测试和CTF完成情况。此类衡量有用但不完整:在运营安全中,每个推理步骤、工具调用、遥测查询和强化请求都会消耗预算。我们通过这种成本-成功率视角,在进攻性的Cybench挑战和防御性的Splunk BOTS v1调查挑战中评估语言模型安全代理。我们不仅报告最佳情况的成功率,还在固定成本水平下比较模型,并按推理花费和工具花费分解性能。结果显示红蓝队任务有不同的扩展模式。进攻性CTF性能随测试时计算量增加而提高,规模化的开放权重模型能接近前沿专有系统并保持成本竞争力。防御性SOC调查则不然:成功更依赖于规范的工具使用、遥测导航和选择性强化,而非仅靠原始推理预算。我们认为安全代理基准应衡量经济效率和与运营的契合度,成本感知的、基于SOC的评估能更清晰地了解哪些模型在当下实际有用,以及防御代理仍需改进之处。我们在此https URL展示了一个包含结果的交互式网站。
英文摘要
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, we compare models at fixed cost levels and decompose performance by inference spend and tool spend. Our results show distinct scalingregimes for red- and blue-team tasks. Offensive CTF performance improves with additional test-time compute, and scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive. Defensive SOC investigation does not scale in the same way: success depends more heavily on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget alone. We argue that security-agent benchmarks should measure economic efficiency and operational fit alongside task success. Cost-aware, SOC-native evaluations provide a clearer picture of which models are practically useful today and where defensive agents still need to improve. We present an interactive website with our results https://evals.frontier.security.