arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10780cs.CR

大到足以突破:追踪LLM渗透测试智能体能力的上升趋势

Big Enough to Break Out: Tracking the Rising Capability of LLM Penetration-Testing Agents

  • University of Virginia(弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Victoria Lovelace, Cameron Berryman, Yuhan You, Suhas Reddy Adavelly, Joel Sadler, Daniel Graham

AI总结:

本研究比较新旧两代LLM渗透测试智能体,发现自主系统能力更强,且限制因素为规划而非记忆,并提出防御者可用的能力追踪方法。

AI中文摘要:

大型语言模型(LLM)智能体正越来越多地被应用于渗透测试,但我们对其能做什么或如何失败仍知之甚少。我们比较了两个基于PentestGPT的系统:一个运行开源权重Kimi K2.5的旧式人在回路系统,以及一个运行Claude Opus 4.8的较新自主系统。在三个公共目标上,自主系统全部解决,包括旧式系统从未完成的两个。旧式系统的结果更令人惊讶。即使在旧式系统未能解决的机器上,它也完成了大约一半的子任务,同时运行在普通大学GPU上且没有提供商护栏。我们可以描述这一趋势但无法解释它,因为模型、框架、自主性和记忆架构都同时变化。其方向仍指向下一个问题:随着这些智能体承担更复杂的任务,什么将限制它们?通常的答案是长时程记忆,即在长攻击链中丢失对早期发现的访问。我们通过向两个系统添加覆盖记忆层来测试这一点,但两者均未改善结果。在我们能审查的旧式停滞运行中,限制因素似乎是规划和承诺而非记忆丢失:智能体持有前进路线的证据却从未将其转化为具体的利用假设,这可能表明进攻能力将随着智能体规划能力的提升而进步,而非依赖于更好的记忆。追踪这种能力的相同子任务评分可供防御者使用,他们可以在其上升时进行测量,而不是等到在实战中遭遇。

英文摘要:

Large language model (LLM) agents are increasingly applied to penetration testing, but we still know little about what they can do or how they fail. We compare two PentestGPT-based systems: a legacy human-in-the-loop system running the open-weight Kimi K2.5, and a newer autonomous system running Claude Opus 4.8. Across three public targets, the autonomous system solves all three, including the two the legacy system never finishes. The legacy result is the more surprising of the two. Even on the machines the legacy system fails to solve, it completes about half the subtasks, while running on ordinary university GPUs with no provider guardrails. We can describe the trend but not explain it, since model, harness, autonomy, and memory architecture all change together. Its direction still points to the next question: what will limit these agents as they take on more complex tasks? The usual answer is long-horizon memory, the loss of access to earlier findings during long attack chains. We test it by adding a coverage-memory layer to both systems, and neither improves outcomes. In the legacy stalled runs we could review, the limiting factor appeared to be planning and commitment rather than lost memory: agents held the evidence for a route forward and never turned it into a concrete exploitation hypothesis, which may suggest that offensive capability will advance with agents' ability to plan rather than with better memory. The same subtask scoring that tracks this capability is available to defenders, who can measure it as it rises instead of waiting to meet it in the field.

补充信息

↑