AI 中文总结
该研究针对自主企业网络防御智能体AI的环境数据缺口,提出Open Security Benchmark框架,通过冻结环境等设计实现端到端评估,还规划了框架的扩展方向。
AI 中文摘要
企业正朝着自主网络防御方向发展:智能体AI可构建企业安全态势的情境感知,并基于此进行评估、决策与行动。这依赖于对企业安全态势的整体视图,即安全态势管理所整合的身份、云与基础设施、数据、应用及其配置的持续跨厂商图景。当智能体承担这项工作时,关键不在于智能体能否给出答案,而在于是否应信任它给出的答案——该领域目前尚无法回答这个问题。真实企业环境是私有的、跨厂商的且深度关联的,且没有任何一个环境作为公开、可查询的目标被暴露出来,用于对这类智能体进行端到端评估,我们将此称为环境数据缺口。我们提出开放安全基准(Open Security Benchmark, OSB),这是一个用于对智能体AI在该工作上进行基准测试的框架。OSB提供了一个精心策划的企业环境——安全态势的冻结整体视图——并通过两种模式评估态势调查:基于关系快照的文本转SQL,以及在同一环境的已部署实例上通过各厂商的原生API进行评估。冻结环境将目标状态固定为不可变快照,并将答案锚定到闭式真值。OSB由五个组件构成:数据层、任务与评估集层、多维评分层、最小可审计工具,以及“自带路径”功能,该功能可从同一底层实现公开比较与私有租户评估。我们用两个身份安全包和一系列覆盖多种规模的合成组织环境数据集实例化该框架,并规划其向更多态势子域、调查模式以及从评估到修复的防御阶段的扩展。
英文摘要
Enterprises are moving toward autonomous cyber defense: agentic AI that builds situational awareness of an organization's security state and reasons from it to assessments, decisions, and actions. This rests on a holistic view of the enterprise's security state, the continuous, cross-vendor picture of identities, cloud and infrastructure, data, applications, and their configurations that security posture management assembles. As agents take on this work, what matters is not whether an agent can produce an answer but whether it should be trusted to. The field cannot yet answer this question. Real enterprise environments are private, cross-vendor, and deeply correlated, and none is exposed publicly as a shared, queryable target for evaluating such agents end to end. We call this the environment data gap. We present Open Security Benchmark (OSB), a framework that benchmarks agentic AI on this work. OSB surfaces a curated enterprise environment - a frozen, holistic view of the security state - and evaluates posture investigation across two modalities: text-to-SQL over a relational snapshot and each vendor's native API over a served instance of the same environment. Freezing the environment pins the target state as an immutable snapshot and anchors answers to a closed-form ground truth. OSB is built from five components: a data layer, a task and evaluation-set layer, a multi-dimensional scoring layer, a minimal auditable harness, and a bring-your-own path that serves public comparison and private tenant evaluation from one substrate. We instantiate the framework with two identity-security packs and a family of synthetic-organization environment datasets spanning multiple scales, and chart its extension to further posture subdomains, investigation modalities, and defense stages from assessment toward remediation.
Comments14 pages, 3 tables. Code: https://github.com/OpenSecurityAI/osb ; datasets: https://huggingface.co/OpenSecurityAI