PatchRisk:预测开源依赖网络中的未来漏洞暴露
PatchRisk: Forecasting Future Vulnerability Exposure in Open-Source Dependency Networks
浏览论文内容
中文总结 AI 辅助
针对开源依赖网络中的未来传递性漏洞暴露预测问题,提出防泄漏基准PatchRisk,并验证HistoryGraph特征在时间测试下显著提升AUPRC。
中文摘要 AI 辅助
开源软件生态系统是网络规模的依赖网络。一个下游软件包可能面临安全风险,并非因为其自身源代码发生变化,而是因为其某个传递依赖项后来收到了漏洞公告。现有的漏洞检测工作通常关注代码当前是否存在漏洞,或已知的易受攻击依赖项是否已经存在。我们研究一个不同的问题:未来的传递性漏洞暴露。给定发布时观察到的软件包-版本依赖图,任务是预测在未来的时间范围内,是否有任何非根依赖项会收到漏洞公告。这个问题对于Web智能和软件供应链安全非常重要,因为它支持在未来的暴露可见之前进行主动的依赖分类。这个问题也容易被错误评估:当前易受攻击的依赖项可能泄露标签,而同一根软件包的不同版本可能在训练集和测试集之间造成软件包级别的记忆。因此,我们构建了PatchRisk,一个基于开源漏洞公告和this http URL依赖图(针对npm和PyPI)的防泄漏基准。该基准使用过滤感知标签、软件包不相交评估、时间测试以及嵌套的1K、3K、5K和10K采样规模。最大的清理设置包含9,007个根软件包-版本图,涵盖4,157个根软件包。我们评估了三个特征族:TimeOnly、GraphStruct和HistoryGraph。在10K时间测试基准上,对于90天预测,HistoryGraph将AUPRC从最强的TimeOnly基线的0.351提高到0.640;对于365天预测,从0.473提高到0.813。这种改进在更小的规模和软件包组洗牌鲁棒性测试中保持稳定。
英文摘要
Open-source software ecosystems are web-scale dependency networks. A downstream package can become exposed to security risk not because its own source code changes, but because one of its transitive dependencies later receives a vulnerability advisory. Existing vulnerability-detection work often focuses on whether code is currently vulnerable or whether a known vulnerable dependency is already present. We study a different problem: future transitive vulnerability exposure. Given a package-version dependency graph observed at release time, the task is to predict whether any non-root dependency will receive a vulnerability advisory within a future horizon. This problem is important for Web intelligence and software supply-chain security because it supports proactive dependency triage before future exposure is visible. It is also easy to evaluate incorrectly: current vulnerable dependencies can leak the label, and different versions of the same root package can create package-level memorization across train and test splits. We therefore construct PatchRisk, a leakage-aware benchmark from Open Source Vulnerability advisories and deps.dev dependency graphs for npm and PyPI. The benchmark uses filtration-aware labels, package-disjoint evaluation, temporal testing, and nested 1K, 3K, 5K, and 10K sampling scales. The largest cleaned setting contains 9,007 root package-version graphs spanning 4,157 root packages. We evaluate three feature families: TimeOnly, GraphStruct, and HistoryGraph. On the 10K temporal-test benchmark, HistoryGraph improves AUPRC over the strongest TimeOnly baseline from 0.351 to 0.640 for 90-day forecasting and from 0.473 to 0.813 for 365-day forecasting. The improvement remains stable across smaller scales and package-group shuffle robustness tests.
发表机构
- Augustana College(奥古纳学院)
机构由 AI 辅助整理,请以论文原文为准。