函数级漏洞评分衡量的是标记率而非模型:配对基准上的协议效应
A Function-Level Vulnerability Score Measures Flag Rate More Than the Model: Protocol Effects on Paired Benchmarks
AI总结:
本研究通过配对基准实验,发现函数级漏洞评分主要反映模型的标记率而非检测能力,评估协议的选择对得分影响显著,且判定常由函数共有文本决定。
AI中文摘要:
语言模型越来越多地被评估为漏洞检测器,类似模型在不同论文中报告的得分差异很大。我们测量了在模型输出保持不变的情况下,评估协议对差异的贡献程度。在配对测试中,模型必须标记易受攻击的函数,并在修复提交后清除其版本。我们逐一改变已发表评估中三个不同的选择:指标、判定提取和输出预算。我们在五个已发布的配对基准和一个为本工作汇集的数据集上,使用同一协议评估了七个前沿模型和大型开放模型,并在汇集的数据集上评估了61个参数量从1.5B到36B的开放模型。函数级F1分数与模型同时标记配对中两个函数的频率高度相关(在42种组合上Spearman相关系数为+0.86),而与配对级正确性几乎无关(+0.16)。在配对分数上,提取在中间值上改变模型数量+0.001,预算改变+0.02且区间包含零,而模型改变基准数量最多达0.18,基准改变模型数量最多达0.16;因此,函数级分数衡量的是标记率而非模型。对于68个模型中的37个,正确配对与反转配对之间的差异在其95%置信区间内包含零,这是以固定比率标记每个函数的零模型的数值,而双标记率和双清除率在中间值上超过零模型0.055,且对于68个模型中的64个,配对中两个函数比独立性预测更常收到一个答案:判定由两个函数共有的文本决定。在长度匹配的配对中,基于激活的线性探针通过配对内排序分离0.78,而tf-idf基线为0.64,长度为0.5;生成的判定在六个模型中的三个上接近随机水平,在其余三个上为0.55至0.57,而提示逻辑回归在所有六个模型上均处于随机水平。
英文摘要:
Language models are increasingly evaluated as vulnerability detectors, and scores reported for similar models differ widely between papers. We measured how much of that difference evaluation protocol accounts for, with model outputs held fixed. In a paired test, a model must flag a vulnerable function and clear its version after a fixing commit. Three choices that published evaluations make differently were varied one at a time: metric, verdict extraction and output budget. Seven frontier and large open models were evaluated on five released pair benchmarks and a set pooled for this work under one protocol, and 61 open models of 1.5B to 36B parameters on the pooled set. Function-level F1 follows how often a model flags both functions of a pair (Spearman $+0.86$ over 42 combinations) and is nearly unrelated to pair-level correctness ($+0.16$). On the pair score, extraction changes a model's number by $+0.001$ at the median and budget by $+0.02$ with an interval through zero, whereas the model changes a benchmark's number by up to 0.18 and the benchmark a model's by up to 0.16; a function-level score therefore measures flag rate more than model. For 37 of 68 models the difference between correct and reversed pairs is within its 95% interval of zero, the value for a null model that flags each function at a fixed rate, while both-flagged and both-cleared rates exceed that null by 0.055 on median, and for 64 of 68 both functions of a pair receive one answer more often than independence predicts: verdicts are determined by the text common to both functions. On length-matched pairs a linear probe on activations separates 0.78 by within-pair ranking, against 0.64 for a tf-idf baseline and 0.5 for length; the generated verdict is near chance for three of six models and at 0.55 to 0.57 for the other three, and a prompted logit is at chance for all six.