从评估到优化:用于Python中CWE预测的层次感知训练信号
From Evaluation to Optimisation: Hierarchy-Aware Training Signals for CWE Prediction in Python
浏览论文内容
中文总结 AI 辅助
研究Python中CWE预测,比较监督微调、双头分类损失、强化学习三种机制,发现监督方法在分布转移下不佳,GRPO成功,最佳策略降低了Qwen2.5-Coder-7B在SVEN数据集上的累积ALPHA惩罚,表明分层惩罚作训练信号的价值取决于传递直接性。
中文摘要 AI 辅助
原始的ALPHA基准引入了一种分类感知惩罚来评估Python中CWE级别的漏洞预测,并提出该惩罚理论上也可作为训练信号。本文对此进行了验证。比较了三种传递机制:监督微调、双头分类损失和基于归一化惩罚的密集奖励的强化学习。发现监督方法在分布转移下始终低于零样本基线,而GRPO成功。最佳策略在贪婪解码下将Qwen2.5-Coder-7B在安全强化和对抗测试(SVEN)数据集上的累积ALPHA惩罚降低了27.9%,在采样解码下降低了25.5%(p = 0.005, Welch检验),与4.5倍大的零样本教师达到统计等效。得出结论,分层惩罚作为训练信号的价值很大程度上取决于其传递的直接性。
英文摘要
The original ALPHA benchmark introduced a taxonomy-aware penalty for evaluating CWE-level vulnerability prediction in Python and proposed that the penalty could theoretically also serve as a training signal. This paper provides that validation. We compare three delivery mechanisms: supervised fine-tuning, a dual-head classification loss, and reinforcement learning with a dense reward derived from the normalised penalty. We find that supervised approaches consistently regress below the zero-shot baseline under distribution shift, while GRPO succeeds. Our best policy reduces the cumulative ALPHA penalty of Qwen2.5-Coder-7B on Security Hardening and Adversarial Testing (SVEN) dataset by 27.9% under greedy decoding, and by 25.5% under sampled decoding(p = 0.005, Welch's t-test), reaching statistical parity with its 4.5x larger zero-shot teacher. We conclude that the value of a hierarchical penalty as a training signal depends largely on the directness of its delivery.