KaliBench:面向Kali Linux网络安全工具使用的细粒度基准,具有免运行时可验证奖励
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
浏览论文内容
中文总结 AI 辅助
KaliBench是一个细粒度基准,用于评估LLM在Kali Linux上将自然语言转换为CLI命令的能力,通过多阶段验证实现免运行时可验证奖励,实验表明8B模型经训练后可媲美685B MoE模型。
中文摘要 AI 辅助
大型语言模型(LLM)越来越多地被应用于网络安全工作流程中,在这些流程中,它们需要将分析师的意图转化为工具调用。然而,现有的评估侧重于基于知识的评估或端到端的智能体任务,并未直接衡量LLM为真实网络安全工具生成可执行命令的能力。这一差距至关重要,因为网络安全操作依赖于严格的命令行接口(CLI),其中微小的语法错误、错误的标志-值绑定或参数顺序错乱都可能导致执行无效。我们引入了KaliBench,一个面向Kali Linux上自然语言到CLI转换的细粒度基准和数据集,包含8,504个查询-命令对,涵盖1,642个工具,跨越23个能力维度和5个安全阶段。KaliBench通过基于手稿的流程构建,具有确定性规范化和别名感知评估,能够对工具选择和参数构建进行精确且可重复的评估。为了确保语义正确性和实际可执行性,我们开发了一个多阶段验证流程,结合了基于LLM的验证、沙盒终端执行和人工参与的细化。基于这些细粒度的确定性信号,KaliBench进一步支持用于训练的免运行时可验证奖励。在三种评估模式和24种通用及安全专注的开源权重模型配置中,没有任何开源权重模型在无限制设置下超过42%的精确命令准确率,这突显了在没有显式工具提示的情况下准确使用基于CLI的网络安全工具的难度。我们进一步表明,基于KaliBench的可验证奖励进行监督微调和强化学习,显著提升了一个8B模型,并达到了与685B MoE模型相当的性能。
英文摘要
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.