arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21471cs.CR

关键基础设施防御的开源情报(OSINT)来源有效性的病例对照测量研究

A Case-Control Measurement Study of OSINT Source Effectiveness for Critical Infrastructure Defense

Ekrem E. Emeksiz, Jeel Piyushkumar Khatiwala, Divyangkumar Patel, Weifeng Xu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过病例对照方法分析2010-2024年54起关键基础设施网络攻击及空案例,将OSINT来源分为三类,发现2类来源覆盖92.6%攻击,k=3贪心组合表现优于随机子集,相关数据与规则已公开。

中文摘要 AI 辅助

关键基础设施(CI)的防御者订阅了大量公共开源情报(OSINT)订阅源,但缺乏哪些订阅源能真正预警攻击的实证依据,本研究填补了这一空白。研究针对2010年至2024年间54起经确认的CI网络攻击,涵盖12个命名CI领域及跨领域类别(见第四节合并规则),并搭配来自同一来源空间的12个空对照漏洞案例,对达到最低量阈值的10类公共OSINT来源,审计其攻击覆盖度、空案例污染度及信号提前时间。来源可清晰分为三类操作上不同的任务轮廓(合并Fisher精确检验p值=3.4×10^-8):前导类(6类,覆盖度≥5%时空案例触发数为0)、披露暴露类(3类,空案例污染度≥攻击覆盖度),以及1类覆盖范围广的混合两类轮廓的来源,其语料库内精度达91.3%。该精度侧分类在2019年时间划分及美国与非美国地理划分中均保持稳定。2类来源(1类广覆盖、1类前导)覆盖92.6%的语料库攻击;3类覆盖96.3%。k=3的贪心组合比平均随机3来源子集表现好39.8个百分点。多个被广泛视为工业控制系统防御标准的来源类,按操作任务属于披露暴露类而非质量问题。尽管存在共同的排名第一的来源,但按领域、行为者及司法管辖区划分的组合排名顺序存在差异。本研究的语料库、关联协议及分类规则已公开。

英文摘要

Defenders of critical infrastructure (CI) subscribe to many public open-source intelligence (OSINT) feeds without an empirical basis for which feeds actually precede attacks. We provide one. Across 54 confirmed CI cyberattacks from 2010 through 2024 spanning twelve named CI sectors plus a cross-sector category (consolidation rules in Section IV), paired with 12 null-control vulnerability cases drawn from the same source space, we audit per-source attack coverage, null-case contamination, and signal lead time for ten public OSINT source classes that meet a minimum-volume threshold. Sources separate cleanly into three operationally distinct mission profiles (pooled Fisher exact p = 3.4x10^-8): precursor (six classes with zero observed null firings at coverage at or above 5%), disclosure-exposure (three classes whose null contamination meets or exceeds attack coverage), and one large broad-coverage class that mixes the two profiles but retains 91.3% within-corpus precision. The precision-side classification is stable across a 2019 temporal partition and across a US-versus-non-US geographic partition. Two sources, one broad-coverage and one precursor, cover 92.6% of corpus attacks; three cover 96.3%. The greedy portfolio at k = 3 outperforms the mean random three-source subset by 39.8 percentage points. Several source classes widely treated as canonical for industrial control system defense fall into the disclosure-exposure profile by operational mission, not by quality. Per-sector, per-actor, and per-jurisdiction portfolios diverge in rank order despite a shared rank-one source. The corpus, linkage protocol, and classification rules are released.

补充信息

↑