arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JevVibe:高效的分类引导安全代码生成

JevVibe: Efficient Classification-Guided Secure Code Generation

Arshak Rezvani, Sasha Behrouzi, Ahmad-Reza Sadeghi

arXiv 2609.34963首次发表:更新:

发表机构

Technical University of Darmstadt(达姆施塔特工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

JevVibe利用Jev决策模型高效分类CWE弱点,引导修复代理提升生成代码安全性,将安全通过率从63.5%提升至70.7%。

AI 中文摘要

大型语言模型能够生成功能正确但仍包含安全弱点的代码,这促使了修复管道的出现,该管道首先诊断弱点类型,然后再决定如何修复。通用弱点枚举(CWE)为此类诊断提供了标准化词汇表,但要求自回归语言模型生成CWE标签并从响应中提取该标签,会引发关于输出有效性、速度、成本以及准确性的问题。我们评估了Jev,一种决策模型,它直接从声明的候选集中选择并返回每个候选的概率,与六个开放权重自回归模型和一个前沿专有模型GPT-5.6-Sol进行比较,在包含1,916个CyberSecEval基准示例的受控50路CWE分类任务上进行测试。Jev在每项分类和排名指标上均优于所有六个开放权重基线,而与GPT-5.6-Sol的比较取决于指标:GPT-5.6-Sol在Top-1准确率和Macro-F1上更高,而Jev在Top-3和Top-5准确率上更高,且MRR几乎相同,同时中位API延迟低6.27倍,估计API成本低55.9倍。我们进一步构建了JevVibe,一个诊断引导的修复代理,使用预测的CWE标签来修复由Qwen2.5-Coder-32B-Instruct生成的代码。借助Jev提供诊断,该代理将检测器测量的安全通过率从修复前的63.5%提高到70.7%,而LLM引导的修复为66.1%。这些结果表明,JevVibe在提高生成代码的安全性方面是有效的,Jev提供了可靠且高效的CWE分类。

英文摘要

Large language models can generate functionally correct code that still contains security weaknesses, motivating repair pipelines that first diagnose a weakness type before deciding how to fix it. The Common Weakness Enumeration (CWE) provides a standardized vocabulary for such diagnoses, but asking an autoregressive language model to generate a CWE label and extracting it from the response raises questions about output validity, speed, and cost, as well as accuracy. We evaluate Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, GPT-5.6-Sol, on a controlled 50-way CWE classification task over 1,916 CyberSecEval benchmark examples. Jev outperforms all six open-weight baselines on every classification and ranking metric, while its comparison with GPT-5.6-Sol depends on the metric: GPT-5.6-Sol achieves higher Top-1 accuracy and Macro-F1, whereas Jev achieves higher Top-3 and Top-5 accuracy and a nearly identical MRR, at $6.27\times$ lower median API latency and $55.9\times$ lower estimated API cost. We further build JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code generated by Qwen2.5-Coder-32B-Instruct. With Jev providing the diagnosis, the agent increases the detector-measured security pass rate from 63.5% before repair to 70.7%, compared with 66.1% for LLM-guided repair. These results show that JevVibe is effective at improving the security of generated code, with Jev providing reliable and efficient CWE classification.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑