发表机构
Purdue University; Red Hat(普渡大学; 红帽公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VEX-Bench是首个评估LLM智能体判断软件供应链漏洞可利用性的基准,含75个真实案例,覆盖Python、Java和Go,实验表明现有模型在细粒度理由分类上仍面临挑战。
AI 中文摘要
软件供应链因其对复杂而脆弱的依赖关系的依赖,已成为日益暴露的攻击面。现有的防御措施(如GitHub Dependabot)通常会引发许多误报,因为其粗粒度的匹配无法确定易受攻击的依赖项是否实际可利用。安全分析师通常需要花费大量时间逐案评估漏洞的可利用性。鉴于LLM智能体在编码和网络安全方面的高级能力,近期已成为该任务的有前景的候选者,但目前尚无基准对其进行评估。先前的基准针对零日漏洞设置,智能体需要检测并利用先前未知的漏洞。相比之下,软件供应链安全关注的是上游依赖中的已知漏洞如何影响下游项目。这要求智能体跨仓库进行推理,并确定上游漏洞在下游项目中是否可利用。为解决这一空白,我们引入了VEX-Bench,这是首个评估LLM智能体评估软件供应链漏洞可利用性能力的基准。它包含75个从GitHub挖掘并由安全专家标注的真实案例,涵盖Python、Java和Go。我们评估了三种智能体框架下的九个模型。虽然GPT-5.5和Claude Opus 4.6在二元漏洞状态分类上达到了约80%的F1分数,但只有GPT-5.5在细粒度理由分类上超过了70%的宏F1分数。这一差距凸显了从二元可利用性评估转向识别细粒度可利用性理由的挑战。代码和数据:此https URL。
英文摘要
The software supply chain has become an increasingly exposed attack surface because of its reliance on intricate yet fragile dependencies. Existing defenses such as GitHub Dependabot often raise many false alerts because their coarse-grained matching cannot determine whether a vulnerable dependency is actually exploitable. Security analysts typically spend substantial time assessing vulnerability exploitability case by case. Recent LLM agents have emerged as promising candidates for this task given their advanced capabilities in coding and cybersecurity, yet no existing benchmark evaluates them on it. Prior benchmarks target zero-day settings, where agents detect and exploit previously unknown vulnerabilities. In contrast, software supply chain security focuses on how known vulnerabilities in upstream dependencies affect downstream projects. This requires agents to reason across repositories and determine whether an upstream vulnerability is exploitable in the downstream project. To address this gap, we introduce VEX-Bench, the first benchmark for evaluating LLM agents' ability to assess the exploitability of software supply chain vulnerabilities. It contains 75 real-world cases mined from GitHub and labeled by security experts, covering Python, Java, and Go. We evaluate nine models across three agent harnesses. While GPT-5.5 and Claude Opus 4.6 reach approximately 80% F1 on binary vulnerability-status classification, only GPT-5.5 surpasses 70% macro-F1 on fine-grained justification classification. This gap highlights the challenge of moving beyond binary exploitability assessment to identifying fine-grained exploitability reasons. Code and data: https://github.com/steven1518/vex-bench
CommentsAccepted to EMNLP 2026