发表机构
Independent Researcher --- AI
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对自主渗透测试中LLM误判问题,提出系统一决策模型(JEV、Laya)支持裁决、校准、剪枝等决策,并设计Rave模型以提升框架可靠性。
AI 中文摘要
自主渗透测试框架使用大型语言模型(LLM)进行侦察、利用和报告,但往往依赖这些相同的模型来确认发现、评定严重性并选择代理。这可能导致误报、严重性夸大和计算资源浪费。我们研究了系统一决策模型(即返回类型化、校准判定结果的轻量级非生成式分类器)如何支持这些决策。我们做出五项贡献。首先,我们定义了四个决策点:发现裁决、严重性重新校准、代理剪枝和确认循环。其次,我们展示了一项探索性 NeuroSploit 案例研究,比较了使用 TypeSafe 系统一(JEV)的一次运行与未使用该模型的一次运行,针对包含 13 个漏洞的 Web 目标。严重性分布、运行时间和按暴露数据类型评分的差异为架构提供了动机,但并未建立统计显著性。第三,我们审查了 JEV、JEV-Ultrafast 和开源 Laya 的已发布规范,而不假设其他基准的结果可迁移至渗透测试。第四,我们讨论了 RLHF、RLAIF、RLCD 和 RLHV 作为训练方法及其对安全决策信任的影响。最后,我们提出了 Rave(一种领域自适应的系统一模型),并概述了其训练数据、评估协议以及对框架保障的潜在影响。
英文摘要
Autonomous penetration-testing harnesses use large language models (LLMs) for reconnaissance, exploitation, and reporting, but often rely on those same models to confirm findings, grade severity, and select agents. This can lead to false positives, inflated severity, and wasted compute. We examine how System One decision models, lightweight non-generative classifiers that return typed, calibrated verdicts, can support these decisions. We make five contributions. First, we define four decision points: finding adjudication, severity recalibration, agent pruning, and confirmation loops. Second, we present an exploratory NeuroSploit case study comparing one run with TypeSafe System One (Jev) and one without it against a web target containing 13 vulnerabilities. Differences in severity distribution, runtime, and grading by exposed data type motivate the architecture but do not establish statistical significance. Third, we review published specifications for Jev, Jev-Ultrafast, and the open-source Laya without assuming that results from other benchmarks transfer to penetration testing. Fourth, we discuss RLHF, RLAIF, RLCD, and RLHV as training approaches and their implications for trust in security decisions. Finally, we propose Rave, a domain-adapted System One model, and outline its training data, evaluation protocol, and potential effect on harness assurance.
Comments21 pages, 5 figures, 11 tables, 4 code listings