基于开放权重模型的神经符号漏洞证明生成
Neuro-Symbolic Proof-of-Vulnerability Generation with Open-Weight Models
浏览论文内容
中文总结 AI 辅助
研究人员提出POVGEN神经符号框架,利用开放权重模型实现低成本漏洞证明生成,在基准测试和真实CVE中表现优于模糊测试与符号执行,还发现了有缺陷补丁及未报告漏洞。
中文摘要 AI 辅助
软件漏洞持续存在,但验证漏洞仍很困难:漏洞证明(PoV)需要能触发漏洞行为的具体输入,而公开的触发输入往往无法获取已披露漏洞。现有技术在有效性、可扩展性、成本和可控性上各有取舍,存在互补设计空间。为弥补这些不足,我们提出POVGEN,一种低成本神经符号框架,通过语义聚焦和基于开放权重模型的大语言模型(LLM)引导约束推理实现高性价比的PoV生成。POVGEN首先定位漏洞相关区域(若有补丁信息则利用),接着执行路径敏感可达性分析,最后通过提取约束并在SMT求解器支持下进行LLM引导推理求解来生成PoV。在近期基准测试中,POVGEN成功为78.98%的漏洞生成PoV,表现优于模糊测试(最高50.20%)和符号执行(2.45%);在250个无公开PoV的真实世界通用漏洞枚举(CVE)中,它为74.80%的案例生成有效PoV,且无补丁信息时仍能复现65.1%的案例。微调后的开放权重模型在关键子任务(即核心约束推理步骤)上可媲美前沿商用LLM,且本地运行无需按样本付费的API成本。应用生成的PoV发现了已披露CVE中的6个有缺陷补丁(均已修复)及5个此前未报告的漏洞(其中4个已获开发者确认并修复)。
英文摘要
Software vulnerabilities are persistent, but validating them remains difficult: a Proof-of-Vulnerability (PoV) requires a concrete input that triggers the vulnerable behavior, yet public triggering inputs are often unavailable for disclosed vulnerabilities. Existing techniques make different tradeoffs in effectiveness, scalability, cost, and controllability, leaving room for complementary designs. To complement them, we present POVGEN, a low-cost neuro-symbolic framework that makes PoV generation cost-effective via semantic focusing and LLM-guided constraint reasoning using open-weight models. POVGEN first localizes vulnerability-relevant regions (utilizing patch information if available), then performs path-sensitive reachability analysis, and finally generates PoVs by extracting and solving constraints with LLM-guided reasoning backed by an SMT solver. POVGEN successfully generates PoVs for 78.98% of vulnerabilities in a recent benchmark, outperforming fuzzing (up to 50.20%) and symbolic execution (2.45%). On 250 real-world CVEs without public PoVs, it generates valid PoVs for 74.80% of cases and reproduces 65.1% when without patch information. The fine-tuned open-weight models match frontier commercial LLMs on key sub-tasks (i.e., the core constraint-reasoning steps) while running locally at no per-sample API cost. Applying the generated PoVs revealed six flawed patches in disclosed CVEs (all subsequently fixed) and five previously unreported vulnerabilities (of which four have been confirmed and fixed by the developers).