评估本地部署SLM在网络安全CTF任务中的上下文分割
Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks
浏览论文内容
中文总结 AI 辅助
针对本地部署SLM在CTF任务中因上下文膨胀导致的性能退化,提出上下文分割两级代理框架,在picoCTF上以更优令牌效率解决18.52%标准代理无法完成的任务。
中文摘要 AI 辅助
高性能开源权重的小型语言模型(SLMs)的普及,使得先进网络安全能力的获取变得民主化,这带来了日益升级的风险,因为这些模型在本地部署时可以绕过专有API的护栏。然而,作为自主代理部署的SLMs在处理长周期、探索性任务(如网络安全夺旗(CTF)挑战)时,常常因上下文膨胀和累积工具调用输出导致的认知退化而表现不佳。为了理解和缓解这一网络安全威胁,我们引入了上下文分割(context segmentation),这是一种两级代理框架,将复杂的利用任务划分为可管理、上下文隔离的子问题。在使用内存受限的gemma-4模型对picoCTF数据集进行评估时,我们证明对于E4B模型,我们的策略充当了一种智能搜索,与暴力重试相比,以更优的令牌效率获得了有竞争力的奖励,并成功解决了18.52%的标准代理执行无法完成的任务。代码可在该https URL获取。
英文摘要
The proliferation of highly capable open-weight Small Language Models (SLMs) democratizes access to advanced cybersecurity capabilities, posing an escalating risk as these models can bypass proprietary API guardrails when deployed locally. However, SLMs deployed as autonomous agents often struggle with long-horizon, exploratory tasks like cybersecurity Capture The Flag (CTF) challenges due to context bloat and cognitive degradation from accumulated tool-call outputs. To understand and mitigate this cybersecurity threat, we introduce context segmentation, a two-level agentic framework that divides complex exploitation tasks into manageable, contextually isolated sub-problems. Evaluating on the picoCTF dataset using memory-constrained gemma-4 models, we demonstrate that for the E4B model, our strategy acts as an intelligent search, achieving competitive rewards with superior token efficiency compared to brute-force retries, and successfully solving 18.52% of tasks that standard agentic execution fails to complete. Code is available at https://github.com/9xeb/context-segmentation.
发表机构
- University of Genoa(热那亚大学)
机构由 AI 辅助整理,请以论文原文为准。