ScopeJudge:用于进攻性安全代理的成本感知预执行门控
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents
浏览论文内容
中文总结 AI 辅助
研究进攻性安全代理中预执行门控,引入ScopeJudge数据集,含4897个工具调用。评估八个法官模型在五种转录策略下表现,发现静态策略不足,推荐成本敏感和召回优先操作点,以支持自主安全代理的实时监控和可扩展监督。
中文摘要 AI 辅助
随着大语言模型(LLM)代理承担进攻性安全工作,单个超出范围的工具调用可能会突破客户的参与边界、扰乱生产或使漏洞赏金发现无效。与固定安全策略不同,重要边界在用户请求中声明且必须从意图推断。进攻性安全的对抗性质加剧了这一挑战,同一工具调用是否超出范围取决于所涉及的目标及其运行上下文,固定策略无法事先枚举。我们研究预执行门控,即一个低成本、可信的LLM法官检查强大、可交换代理提出的每个调用,并在其运行前接受或拒绝。我们引入ScopeJudge,它是一个包含4897个工具调用的基准(7.7%超出范围违规),来自旨在诱使代理超出范围的任务的代理轨迹,并由专业渗透测试人员在调用级别标记,具有较高的评分者间一致性(Fleiss kappa = 0.64),设定了F1 = 0.78的专家一致参考点。我们在五种转录策略下评估八个法官模型,改变法官所见上下文量,从仅静态策略到完整原始转录,并绘制由此产生的成本-准确性帕累托前沿。我们发现静态策略在结构上不足以进行范围执行:对用户请求视而不见,法官召回率降至接近零,证实范围存在于请求中且基于请求的监控是必要的。由于错过违规的成本高于虚假拒绝,我们分别报告精度、召回率和F1,并推荐两个操作点:一个成本敏感配置和一个高风险部署的召回优先配置。我们发布ScopeJudge数据集以支持对自主安全代理的实时监控和可扩展监督。
英文摘要
As LLM agents take on offensive security work, a single out-of-scope tool call can breach a client's engagement boundary, disrupt production, or void a bug-bounty finding. Unlike a fixed safety policy, the boundary that matters is declared in the user's request and must be inferred from intent. That challenge is sharpened by the adversarial nature of offensive security: the same tool call is in or out of scope depending not on the action itself but on the target it touches and the context in which it runs, which no fixed policy can enumerate in advance. We study pre-execution gating: a cheap, trusted LLM judge inspects each call proposed by a strong, swappable agent, and accepts or rejects it before it runs. We introduce ScopeJudge, a benchmark of 4,897 tool calls (7.7% scope violations) from agent trajectories on tasks engineered to tempt agents out of scope and labeled at the call level by professional penetration testers, with substantial inter-grader agreement (Fleiss kappa = 0.64) that sets an expert agreement reference point of F1 = 0.78. We evaluate eight judge models under five transcript strategies, varying how much context the judge sees, from the static policy alone to the full raw transcript, and chart the resulting cost-accuracy Pareto frontier. We find that a static policy is structurally insufficient for scope enforcement: blind to the user's request, judge recall collapses to near zero, confirming that scope lives in the request and that request-conditioned monitoring is necessary. Because a missed violation costs more than a spurious rejection, we report precision, recall, and F1 separately and recommend two operating points: a cost-sensitive configuration and a recall-first one for high-stakes deployments. We release the ScopeJudge dataset to support real-time monitoring and scalable oversight of autonomous security agents.