发表机构
Qatar Computing Research Institute (QCRI); HBKU(卡塔尔计算研究所(QCRI); 哈马德·本·哈利法大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对静态基准的缺陷,提出动态框架SSP-Bench,通过多目标优化生成评估实例,在24个模型上揭示静态评估的系统性失败,推动部署相关评估。
AI 中文摘要
大型语言模型(LLM)在安全、安保与隐私(SSP)方面的评估严重依赖静态基准,这些基准存在分数饱和、数据污染和聚合伪影等问题,并且无法捕捉对语言变异的敏感性。因此,在固定测试集上表现良好的模型,在面对语义等价的改写时往往失败。我们引入了SSP-Bench,一个动态基准测试框架,它按需生成评估实例,同时保持领域一致性。该框架通过外部来源确保标签有效性,通过服务特定验证强制范围,并使用多模型引导面板校准难度。基准构建被表述为一个关于难度、可分离性、新颖性和多样性的多目标优化问题。在24个模型和四个SSP服务中,SSP-Bench揭示了静态评估的系统性失败,包括由于构念混合导致的安全排名近乎零相关、安全与过度拒绝之间的强耦合,以及隐藏的族内回归。这些结果表明,静态基准可能歪曲模型行为,从而推动了动态、与部署相关的评估。
英文摘要
Evaluation of large language models (LLMs) for safety, security, and privacy (SSP) relies heavily on static benchmarks, which suffer from score saturation, data contamination, and aggregation artifacts, and fail to capture sensitivity to linguistic variation. As a result, models that perform well on fixed test sets often fail under semantically equivalent rephrasings. We introduce SSP-Bench, a dynamic benchmarking framework that generates evaluation instances on demand while preserving domain consistency. The framework ensures label validity through externally grounded sources, enforces scope via service-specific validation, and calibrates difficulty using a multi-model steering panel. Benchmark construction is formulated as a multi-objective optimization problem over difficulty, separability, novelty, and diversity. Across 24 models and four SSP services, SSP-Bench reveals systematic failures of static evaluation, including near-zero correlation in safety rankings due to construct mixing, strong safety--over-refusal coupling, and hidden within-family regressions. These results show that static benchmarks can misrepresent model behavior, motivating dynamic, deployment-relevant evaluation.