arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基准分数依赖于评估流水线:网络安全大语言模型基准的可靠性审计

Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks

Aymene Berriche, Cathrine Shalby, Mohannad Alhanahnah, Yazan Boshmaf

arXiv 2609.08765首次发表:更新:

发表机构

Qatar Computing Research Institute, HBKU(哈马德·本·哈利法大学卡塔尔计算研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究审计8个网络安全基准与10个大语言模型,发现评估流水线选择可致分数变动超80个百分点并改变排名,提出流水线感知审计作为可靠评估的核心要求。

AI 中文摘要

大语言模型(LLM)基准通常被视为具有稳定分数的固定数据集,然而其结果取决于可配置的评估流水线。我们对10个专有、开放权重及网络安全专用大语言模型在8个网络安全基准上进行了审计。通过将基准建模为测量流水线,我们识别出15种系统性故障模式,并表明单一的流水线选择可使模型分数变化超过80个百分点,且显著改变模型排名。在跨基准层面,两对语义相似的任务对因评估约定不兼容而对相同模型产生不同排名。在标准化流水线选择同时保留任务语义的评估框架下,10个模型中有9个在至少一个基准上排名变动至少3位。这些结果表明,网络安全大语言模型基准分数依赖于流水线,并促使将流水线感知审计作为可靠模型评估的核心要求。

英文摘要

Large language model (LLM) benchmarks are often treated as fixed datasets with stable scores, yet their outcomes depend on configurable evaluation pipelines. We audit eight cybersecurity benchmarks across 10 proprietary, open-weight, and cybersecurity-specialized LLMs. By modeling benchmarks as measurement pipelines, we identify 15 systematic failure modes and show that a single pipeline choice can change a model's score by more than 80 percentage points and substantially alter model rankings. At the cross-benchmark level, two semantically similar task pairs rank the same models differently because of incompatible evaluation conventions. Under an evaluation harness that standardizes pipeline choices while preserving task semantics, nine of 10 models shift by at least three ranks on at least one benchmark. These results show that cybersecurity LLM benchmark scores are pipeline-dependent and motivate pipeline-aware auditing as a core requirement for reliable model evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑