发表机构
Tel Aviv University; University of Massachusetts Amherst(特拉维夫大学; 马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用地缘政治数据和大语言模型预测网络安全事件,结合网络测量提升准确性,优于现有技术,并分析了数据源的有效性。
AI 中文摘要
预测安全事件是一项意义深远的任务,对于制定主动防御措施和网络保险政策至关重要。先前处理此问题的研究主要利用基于网络测量(如协议配置错误)的结构化、人工定义的特征。尽管这些方法取得了令人鼓舞的性能,但基于网络的特征可能无法捕捉与攻击者动机相关的方面。为填补这一空白,我们的工作利用从公开来源挖掘的地缘政治数据——这可能有助于捕捉攻击者的动机——来预测安全事件。具体而言,我们的方法依赖于新闻文章和转录的播客,这些内容被输入到大语言模型中,以自动生成丰富的表示。随后,这些表示被输入到一个分类器,该分类器基于历史事件训练以预测未来事件。我们使用一个大型事件数据集(超过15,700条记录)进行的评估表明,仅使用地缘政治数据就能实现显著的准确性(ROC AUC为71.4%)。值得注意的是,将地缘政治数据与网络测量相结合,其性能优于仅基于网络特征的现有最先进技术(ROC AUC分别为81.3%和77.5%)。我们的分析还有助于揭示地缘政治数据在何时最为有用,以及哪些数据源对准确预测最有帮助。
英文摘要
Predicting security incidents is a profound task critical for informing proactive defensive measures and cyber-insurance policies. Prior work tackling this problem mainly utilized structured, manually defined features based on network measurements (e.g., protocol misconfigurations). Still, despite leading to promising performance, the network-based features may fail to capture aspects related to adversaries' motives. To fill this gap, our work leverages geopolitical data mined from public sources--which may help capture attacker motives--to forecast security incidents. Specifically, our approach relies on news articles and transcribed podcasts that are fed to large language models to automatically produce rich representations. The representations are then fed to a classifier trained to forecast future incidents based on historical ones. Our evaluation with a large incidents dataset (>15,700 records) demonstrates substantial accuracy (71.4% ROC AUC) with geopolitical data alone. Notably, combining geopolitical data and network measurements outperforms the state-of-the-art technique based on network features alone (81.3% vs.77.5% ROC AUC). Our analysis also helps shed light on when geopolitical data is most helpful and the data sources that are most useful for accurate forecasting.