arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基准的基准:评估小型语言模型的自动化安全基准

Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

Nyamtulla Shaik, Fengjun Li, Bo Luo

arXiv 2608.17183首次发表:更新:

发表机构

University of Kansas(堪萨斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过评估5个基准套件在26个开源SLMs上的表现,发现现有以LLM为中心的安全基准无法可靠评估SLMs,模糊性会导致模型排名出现显著变化。

AI 中文摘要

小型语言模型(SLMs)越来越多地部署在资源受限、隐私敏感的场景中,在这些场景中,安全和偏见故障可能会导致安全和社会风险。然而,现有的AI安全/安全/合规基准是为大型语言模型设计的,可能无法可靠地迁移到SLMs。因此,我们提出问题:这些基准能否有效且可靠地评估SLMs?为了回答这个问题,我们通过在统一评判标准下评估五个广泛使用的基准套件在26个开源SLMs上的有效性和鲁棒性,开展了大规模评估,该标准对有害、安全或模糊/不相关的响应分别赋予0、1或0.5分。在所有基准中,模糊判断占主导地位,且与提示复杂度和模型架构相关,这表明以大型语言模型(LLM)为中心的安全基准不足以作为SLM安全评估的独立证据。总体而言,模糊率随词汇密度、输出困惑度和输出长度增加而上升,随词汇复杂度、自一致性以及回复与提示的相似度增加而下降。这揭示了将模型能力与表观安全相混淆的能力-安全混淆。由于模糊性普遍存在,聚合平均分数排行榜在数学上是脆弱的:即使底层输出保持不变,在合理的模糊处理下,模型排名也会发生显著变化。

英文摘要

Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures can cause security and societal risks. However, existing AI safety\slash security\slash compliance benchmarks are designed for large language models that may not transfer reliably to SLMs. We therefore ask: Can these benchmarks effectively and reliably evaluate SLMs? To answer this question, we conduct a large-scale assessment of the effectiveness and robustness of these automated pipelines by evaluating five widely used benchmark suites across 26 open-source SLMs under a unified judging rubric, which assigns a score of 0, 1, or 0.5 to harmful, safe, or ambiguous/irrelevant responses, respectively. Across the benchmarks, ambiguous judgments dominate and correlate with prompt complexity and model architecture, indicating that {\em LLM-centric safety benchmarks are insufficient as standalone evidence for SLM safety assessment}. In general, the ambiguity rate increases with lexical density, output perplexity, and output length and decreases with lexical sophistication, self-coherence, and reply-prompt similarity. This reveals a capability-safety confound that mixes model capability with apparent safety. Since ambiguity is prevalent, aggregate mean-score leaderboards are mathematically brittle: model rankings change significantly under reasonable ambiguity treatments, even when the underlying outputs remain unchanged.

CommentsThis paper is accepted for publication at ESORICS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑