发表机构
Sola Security(索拉安全公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对云安全群体调查任务,评估专用安全大脑 Sola Security Brain 与通用编码代理 Claude Code 的性能,结果显示前者覆盖率提升 79.2%,成本大幅降低,并揭示代理的“采样-泛化”缺陷。
AI 中文摘要
云安全调查主要由群体任务主导:哪些身份可以读取数据存储,多少资源未通过某项控制检查,哪些资产可从另一个账户访问。这些任务需要针对完整清单进行解析,而非针对某个命名对象。对其中一个问题的部分回答并非部分结果,而是不同的结果。通用编码代理现在可以被授予只读云凭证并直接进行调查,这引发了一个问题:专门构建的安全上下文层还能贡献什么。我们评估了 Sola Security Brain——一个安全智能层,其关系型底层离线解析,安全逻辑在查询时针对该底层进行评估——与 Claude Code 在同一实时 AWS 环境中通过只读 CLI 操作进行对比,覆盖 28 个云安全调查任务。答案通过盲评、分层加权、基于依据门控的相对召回率(基于联合声明池)进行评分,并在三次独立评分抽取中取平均。Sola Security Brain 达到 0.693 的覆盖率,而对比为 0.387,差距为 0.306,三次评分抽取的差异为 ±0.018,即相对增益 79.2%。它在 28 个任务中的 25 个上领先于较弱模型层级,每个任务的推理成本低 17.7 倍,每单位覆盖率的成本低 31.6 倍。除总体结果外,我们描述了一种答案层面的模式,称为“采样-泛化”:在回合预算下,实时代理枚举大型群体的一小部分,断言无保留的全称否定,并且仅在答案元数据中披露样本量,而非在答案本身中。在一个任务中,它在检查了四个存储桶族后报告不存在存储桶策略,而该扫描仅采样了约 5,000 个存储桶中的 40 个,且该账户中有 65 个存储桶带有通配符主体读取授权。
英文摘要
Cloud-security investigation is dominated by population tasks: which identities can read a data store, how many resources fail a control, what is reachable from another account. These resolve against a complete inventory, not a named object, so a partial answer to one is not a partial result but a different one. Coding agents can now be given read-only cloud credentials and asked to investigate directly, which raises the question of what a purpose-built security context layer still contributes. We evaluate the Sola Security Brain, a security intelligence layer whose relational substrate is resolved offline and whose security logic is evaluated against it at query time, against two coding agents operating the same live AWS environment through a read-only CLI, Claude Code and OpenAI Codex, over 28 investigation tasks. All three are scored by a blinded, tier-weighted, grounding-gated relative recall over one joint claim pool, so the scores share a denominator. The Sola Security Brain reaches 0.549 +/- 0.012 coverage against 0.340 +/- 0.011 for Claude Code and 0.281 +/- 0.006 for Codex, with the ordering identical in every grading draw. It leads 24 of 28 tasks from the cheapest model tier, at 17.8x and 20.8x lower cost per task than Claude Code and Codex. Beyond the aggregate, we describe an answer-level pattern we term sample-and-generalise: both agents enumerate a fraction of a large population, assert an unhedged universal negative, and disclose the sample size only in answer metadata rather than in the answer. In one task Claude Code reported that no bucket policies exist after checking four bucket families, in a sweep that sampled 40 of roughly 5,000 buckets, in an account where 65 buckets carry a wildcard-principal read grant.
Comments13 pages, 7 tables